What STEP 8 tells us about maintenance, and what it does not
Every withdrawal trial compared full dose against nothing. The clinically interesting comparison — full dose against a reduced one — has not been randomised.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Interruption
The best argument for tapering is behavioural rather than pharmacological, and it deserves to be made on its own terms.
That does not settle the question, because the best argument for a gradual reduction is not pharmacological. It is that the return of appetite is the thing being managed, that it returns over weeks rather than instantly, and that a person given a month of partial appetite suppression in which to establish different eating patterns may be better placed than a person whose appetite returns fully while they are still eating as though it had not. That is a plausible behavioural hypothesis. It has not been tested in a randomised trial and the Journal is not aware of one being planned.
The parent programmes establish the losses from which the withdrawal arms fall: approximately 14.9 per cent at sixty-eight weeks for semaglutide 2.4 mg in adults without diabetes, approximately 20.9 per cent at seventy-two weeks for tirzepatide 15 mg, and — the only continuous two-year comparator anybody has — approximately 15.2 per cent sustained at week 104 with treatment maintained throughout.12
Read together, the withdrawal evidence supports four statements and does not support a fifth. Regain begins promptly after cessation, within weeks rather than months. It proceeds at a decelerating rate, with the steepest portion in the first three to six months. It does not, within twelve months of follow-up, return participants fully to their original baseline; residual reductions of roughly five to ten per cent persist at one year in all three datasets. And continued treatment maintains and usually extends the loss, with the extension diminishing as the plateau is approached.
The statement not supported is that the drugs cause weight regain, or that stopping leaves a person worse off than never having started. Nothing in these datasets shows overshoot above the original baseline at a group level. Every arm that stopped remained below where it began at the end of follow-up.
The Journal makes this point repeatedly because the contrary claim circulates widely and is often accompanied by a mechanistic story about metabolic damage. The withdrawal trials are the direct test of that claim and they do not support it. What they do support is the unremarkable proposition that a treatment for a chronic condition works while it is being taken.
Mean regain trajectories conceal a distribution as wide as the one that characterises the initial response. In the reported withdrawal arms, some participants returned to within a percentage point of their original baseline within a year while others retained most of their loss with no pharmacological support whatever. The interquartile ranges in the published figures are broad, and the trials were not designed to explain them.
Nothing measured at randomisation predicts an individual regain trajectory usefully. Baseline body mass index, the magnitude of the initial loss, age, sex and diabetes status all shift the mean modestly and leave the spread largely intact. The behavioural variables that plausibly matter — what a person was eating and doing during the loss, and whether any of it persisted — were not measured with the granularity required to test them.
This is the same structural gap that runs through the whole field: the mean effects are well characterised and the variance is not. It has a specific practical consequence here. A person deciding whether to stop cannot be given a personal probability of holding their weight, because no such probability has been established, and any source offering one is offering a group mean dressed as a prediction.
That a treatment for a chronic condition stops working when it is stopped is not a finding about the treatment. It is a finding about the condition.
On how the withdrawal trials were receivedThe Journal has asked clinicians in four jurisdictions how they manage maintenance and received a broadly consistent description that appears in no guideline. Reduce by one escalation step once the weight has been stable for a period; hold for eight to twelve weeks, which is long enough for the new exposure to reach steady state and for a trend to become visible; if the weight rises by more than a small threshold, return to the previous step. Some reduce again after a further stable interval; most do not go below the second step.
Two things recommend this approach and neither is evidence. It follows the pharmacokinetics, in that eight to twelve weeks is comfortably longer than the four to five weeks required to reach steady state at the new dose, so the observation is not being made on a still-changing exposure. And it is reversible, which a decision to stop is not in the same easy way.
The Journal reports this as description, not endorsement. It is not a dosing recommendation, no trial supports it, and the appropriate person to design a maintenance strategy is a clinician who knows the patient. We report it because a practice this widespread deserves to be described accurately rather than left to circulate in fragments.
| Reason | Randomised evidence on outcome | Typical notice | Resumption likely? |
|---|---|---|---|
| Protocol-driven withdrawal | Three designs | Planned | Not applicable |
| Reached target weight | None | Planned | Sometimes |
| Intolerable side effects | Discontinuation rates only | Days | Sometimes, lower dose |
| Cost or coverage loss | None | Weeks or none | Often, when coverage returns |
| Supply interruption | None | None | Usually, at reset tolerability |
| Discontinuation rates for adverse events are reported in every pivotal trial; outcomes after discontinuation for the other reasons are not, because the trials did not enrol people who stopped for them. | |||
This section is short because the evidence is. The Journal has searched the trial registries and the published literature for any randomised comparison of an intermittent schedule against a standard weekly schedule for any GLP-1 receptor agonist or dual agonist, at any dose, for any indication. We have found none. We have also found no observational cohort large enough to characterise outcomes on such a schedule with the standard confounders addressed.
What exists is dose-ranging data from the phase 2 programmes, which establishes that lower average exposures produce smaller weight effects, and pharmacokinetic modelling, which establishes what average exposure and what peak-to-trough ratio a given interval would produce. Neither tells you whether a fortnightly schedule maintains weight in somebody who has already lost it, which is the question actually being asked.
An absence of evidence is not evidence of harm and the Journal does not present it as such. It is, however, the entire evidentiary position, and readers encountering confident protocols for intermittent use should know that the confidence is not coming from data. Nothing in this section is advice, and the compounds sold for research use only that appear in some of these protocols are not approved for human use.
The withdrawal question changes shape when the drug was prescribed for something other than weight. In the cardiovascular outcome trial of semaglutide in overweight and obesity without diabetes, the reduction in major adverse cardiovascular events emerged over years of continued treatment, and the trial provides no information about what happens to that benefit on cessation.3 The same applies to the renal outcome data in chronic kidney disease with type 2 diabetes, where the effect on kidney disease progression was measured over a median of several years of treatment.4
There is no reason to expect an outcome benefit that accrues over years to persist after the exposure ends, and no trial has tested it. For a person taking the drug for glycaemic control, stopping has an immediate and measurable consequence in HbA1c over the following three months. For a person taking it for cardiovascular or renal risk, stopping has no measurable short-term consequence at all, which makes the decision harder rather than easier.
This is the situation in which the Journal thinks the withdrawal-trial coverage has done the most damage. Framing discontinuation as a weight question invites a person taking the drug for kidney disease to reason about it in the wrong currency entirely.
There is no withdrawal syndrome from these agents, no dependence, and no pharmacological reason to reduce gradually rather than to stop. A seven-day half-life produces its own taper: concentrations halve within a week and fall to a few per cent within a month regardless of intent. On the pharmacology alone, a planned taper accomplishes nothing that stopping does not.
The behavioural argument is different and better. Appetite returns over weeks. A person whose dose is reduced in steps experiences that return in stages, while continuing to have some pharmacological support, and has a window in which to establish eating patterns that will have to hold without the drug. A person who stops outright experiences the same return without that window. Whether the window produces better outcomes is an empirical question that has not been asked in a trial.
The Journal’s position is that the behavioural argument is worth making on its own terms and worth not dressing in pharmacological clothing. What a taper cannot do is prevent regain, since the withdrawal trials establish that ongoing exposure is what holds the weight. Presenting a taper as a way of stopping without regaining is a claim the evidence does not support in any form.
Exposure falls on a predictable schedule. A week after a final weekly injection roughly half the steady-state concentration remains; at two weeks a quarter; at four weeks a few per cent. Appetite returns along a curve that lags the concentration curve somewhat, and delayed gastric emptying resolves over a similar interval, so the sensation of early fullness that has been governing meal size for months disappears over about a month.
What people report, and the Journal reports it as report rather than as measurement, is that the appetite return is experienced as abrupt rather than gradual — as a threshold being crossed somewhere in the third or fourth week rather than as a smooth ramp. That is consistent with a non-linear relationship between receptor occupancy and the perceived effect, which is what the dose-response data would predict, and it is not consistent with the exposure curve alone.
Glycaemic parameters follow their own timescale. Fasting glucose responds within days to weeks; HbA1c, which reflects the preceding three months with the recent weeks weighted most heavily, will not show the full consequence of stopping for a full quarter. A panel drawn four weeks after cessation will understate the change, and this is a specific and common misreading. In the trial population with type 2 diabetes, where glycaemic control was a co-primary concern, the treatment effect on HbA1c was of the order of one and a half percentage points, which is the scale of what a cessation eventually undoes.5
Every withdrawal trial compared a full dose against nothing. The comparison almost every patient actually faces has never been randomised.
On the maintenance gapEvery trial in this class delivers a behavioural intervention alongside the drug: energy-restriction targets, activity targets, and regular contact with a study team. That contact is itself an intervention of measurable effect, which is why placebo arms in these programmes lose two to three per cent of body weight rather than nothing. Where the behavioural component was deliberately intensified, the placebo arm lost around 5.7 per cent over sixty-eight weeks, which is a useful upper bound on what contact and counselling alone achieved in these populations.6
It matters for the withdrawal question in a way that is usually elided. The semaglutide off-treatment extension withdrew the drug and the lifestyle support together, so its regain figure describes the removal of a package.7 The STEP 4 and SURMOUNT-4 withdrawal arms kept the lifestyle component running, so their regain figures describe the removal of a molecule with support maintained.89 Those are different experiments and the second is the more conservative.
Anybody comparing regain figures across the three should therefore expect the extension to look worse, and it does. The Journal states which withdrawal design a figure comes from every time it quotes one, because the alternative is pooling two different experiments into a single number that describes neither. The same caution applies to the frequent comparison with dietary weight-loss regain, where the behavioural intervention is the whole of the treatment.
| Study and arm | At randomisation | At end of follow-up | Change during follow-up |
|---|---|---|---|
| STEP 4, continued semaglutide | −10.6% | −17.4% | −7.9% |
| STEP 4, switched to placebo | −10.6% | ≈ −5% | +6.9% |
| SURMOUNT-4, continued tirzepatide | −20.9% | −25.3% | −5.5% |
| SURMOUNT-4, switched to placebo | −20.9% | −9.9% | +14.0% |
| STEP 1 extension, former semaglutide | −17.3% at wk 68 | −5.6% at wk 120 | ≈ +11.6% |
| All values are percentage change from original trial baseline, treatment-policy estimand where reported. The STEP 1 extension figure is an off-treatment observation in a subset and is not comparable with the randomised rows above it. | |||
Four things accompany every regain number in these pages. Which withdrawal design it comes from, because an off-treatment extension and a randomised placebo switch are different experiments. Whether the lifestyle intervention continued in the arm being described. What the denominator is — regain as a percentage of body weight, as a percentage of the weight lost, or as a final position relative to original baseline, three quantities that are routinely quoted interchangeably. And the follow-up duration, because the regain curve decelerates and a figure at six months is not a figure at a year.
The third of those is where most of the misreporting happens. A statement that participants regained two-thirds is a proportion of loss; a statement that they regained eleven per cent is a proportion of body weight; a statement that they finished 5.6 per cent below baseline is a final position. All three can describe the same arm and they are not interchangeable.
Where a source we are quoting has not stated its denominator, we say that rather than inferring it. Readers who find a regain figure in these pages without its design and its denominator have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
This is reporting on a body of trial evidence and it is not advice about whether or how to stop taking a medicine. The decision to discontinue an agent prescribed for glycaemic control, cardiovascular risk or kidney disease is materially different from the decision to discontinue one prescribed for weight, and in every case it belongs with a clinician who has seen the person and knows why the drug was started.
Two further notes. Compounds sold for research use only are not approved for human use in any jurisdiction, and nothing here should be read as guidance about using them or about stopping their use. And where this piece describes what clinicians report doing about maintenance dosing, that is description of practice and not a schedule anybody should adopt from a magazine.
The Journal takes correspondence on this subject at letters@compoundjournal.com and factual challenges at standards@compoundjournal.com. Letters describing a personal experience of stopping are read with attention and are published, where they are published, as accounts rather than as evidence — a distinction this department tries hard to preserve in both directions.
What the shortage years demonstrated, at a scale no trial will ever match, is that this class is now being stopped and restarted routinely by circumstance rather than by decision. That is the discontinuation that actually happens, and there is no literature on it at all. The Journal regards documenting it as reporting rather than research, and will keep doing so.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
I take this for kidney disease, not for weight. Every piece of writing I encounter about stopping is about the weight coming back. It has taken me a year to find anybody willing to say plainly that the renal benefit accrued over years of treatment and nobody has tested what happens if I stop.
— K. Rautio, Tampere
The claim that stopping does not leave you worse off than baseline is a group-level claim about trial arms. Individuals can and do overshoot. Your phrasing invites readers to conclude otherwise.
— M. Ferrari, Trieste
Correct, and the distinction matters. We have added a clause: no arm overshot at a group level, which is not the same as no participant overshooting. The trials do not report individual overshoot rates and we have not found them published anywhere.
Three months after stopping, my HbA1c had barely moved and I concluded I had got away with it. Six months after stopping, it was back where it started. Your point about the lag is the single most useful sentence I have read on this subject.
— C. Wilcoxson, Des Moines, IA
Every withdrawal trial compared full dose against nothing. The clinically interesting comparison — full dose against a reduced one — has not been randomised.
Half-life, accumulation ratio and time to steady state are three separate quantities, and confusing them produces most of the bad advice in circulation.
Nausea and gastric delay attenuate over weeks at an unchanged dose. That single physiological fact is the entire justification for holding.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
The rule that a quarter of weight lost is lean tissue has been in textbooks for decades and does not survive close reading.
Almost every misreading of a laboratory panel is a misunderstanding of what a reference interval is and how much a result has to move before the movement means anything.