Continued treatment against continued placebo: the comparison that settled it
The evidence on stopping is better than the evidence on almost anything else in this field, because somebody deliberately randomised it.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Escalation
The published ladder exists because a protocol needed a single number. Practice has never followed it exactly, and the regulatory file never assumed it would.
There is a particular kind of error that comes from reading a dose table as an instruction rather than a default. A person who reaches week seventeen, feels unwell, and escalates anyway because the calendar says so is following the document and ignoring the drug. A person who reaches an adequate response at an intermediate dose and escalates anyway because the top of the ladder is the top of the ladder is doing something the trials give no reason to do. Both errors have the same root: mistaking the shape of a protocol for the shape of a treatment.
A titration schedule in a product label is the residue of a protocol decision. Somebody designing a phase 3 trial had to specify how participants would arrive at the target dose, because a trial cannot be run on clinical judgement without becoming uninterpretable. The specification was chosen to be tolerable enough that dropout would not wreck the analysis, and brisk enough that the primary endpoint would arrive within the planned duration. It was then submitted, reviewed, approved and printed.
Nothing in that sequence involves comparing the chosen schedule against a plausible alternative. The regulatory question is whether the product as studied is safe and effective, not whether the escalation pattern used to study it was optimal. Assessment reports for the major incretin products discuss dose selection at length and escalation interval selection barely at all.
This is not a criticism of the regulators, who answered the question they are asked. It is a caution against a specific inference: that because a schedule appears in an approved label, it has been shown to be better than a slower or faster one. It has not. The Journal treats the printed ladder as a well-reasoned default with a pharmacokinetic justification, which is what it is.
For weight management, the approved escalation runs 0.25 mg weekly for four weeks, then 0.5 mg, then 1.0 mg, then 1.7 mg, reaching 2.4 mg at week seventeen. For glycaemic indications the ladder is shorter and the maximum lower: 0.25 mg, then 0.5 mg, then 1.0 mg, with a 2.0 mg option added later on the strength of a dedicated dose-comparison study.
Two features are worth noticing. The starting dose is explicitly sub-therapeutic — 0.25 mg is a tolerability rung, not a treatment dose, and describing it as a low dose rather than an initiation dose causes real confusion. And the ratios narrow as the ladder rises: two doublings, then a 1.7-fold step, then a 1.41-fold step.
The 2.4 mg dose was selected on the basis of the phase 2 dose-ranging programme and carried into the STEP trials, where it produced a mean weight reduction of about fifteen per cent at sixty-eight weeks against roughly two and a half per cent on placebo.1 That is the number the ladder exists to reach, and the ladder itself was never the subject of the trial.
The class has two dose-response curves running in parallel, and only one of them flattens.
On why the ceiling existsTirzepatide begins at 2.5 mg weekly for four weeks, moves to 5 mg, and thereafter increases in 2.5 mg increments at intervals of not less than four weeks, to a maximum of 15 mg. The structural difference from semaglutide is important: after the first doubling the increments are fixed in absolute terms, which means the ratio falls steadily — 1.5-fold, then 1.33, then 1.25, then 1.20.
The practical consequence is that the upper half of the tirzepatide ladder is unusually gentle in proportional terms, and the first step from 2.5 mg to 5 mg is by some distance the most demanding thing the schedule asks. Clinicians we spoke to described the 2.5-to-5 transition as the point at which most early attrition occurs, which is what the ratios predict.
The label also states, in language that repays attention, that 5 mg is a therapeutic dose in its own right and that escalation beyond it should reflect response and tolerability. That is a materially different instruction from a ladder with a fixed destination, and it is closer to how the drug is actually used.2
| Programme | Molecule | Dose | Duration | Mean weight change |
|---|---|---|---|---|
| STEP 1 | Semaglutide | 2.4 mg weekly | 68 weeks | −14.9% |
| STEP 1 | Placebo | — | 68 weeks | −2.4% |
| STEP 5 | Semaglutide | 2.4 mg weekly | 104 weeks | −15.2% |
| SURMOUNT-1 | Tirzepatide | 5 mg weekly | 72 weeks | −15.0% |
| SURMOUNT-1 | Tirzepatide | 10 mg weekly | 72 weeks | −19.5% |
| SURMOUNT-1 | Tirzepatide | 15 mg weekly | 72 weeks | −20.9% |
| SURMOUNT-1 | Placebo | — | 72 weeks | −3.1% |
| Treatment-policy estimand where reported. Figures are means from the primary publications and are not comparable across programmes, which differed in population, duration and analysis. | ||||
The pivotal protocols in this class permitted escalation to be delayed. In the semaglutide obesity programme, participants unable to tolerate a dose increase could remain at the previous dose and attempt escalation later; the tirzepatide programme contained comparable provisions, with defined windows and a limit on how long a participant could remain below target before being counted as not having reached it.
This matters for how the efficacy figures should be read. The mean weight reductions quoted from STEP and SURMOUNT were produced by populations in which a meaningful minority spent time below their assigned target dose. The trials therefore already contain the effect of flexible titration, and the flexible approach is not a departure from the evidence but part of how the evidence was generated.
What the labels carry forward is the ladder. What they largely omit is the permission. A prescriber who reads only the label sees a fixed calendar; a prescriber who reads the trial documentation sees a calendar with a documented escape valve. The Journal has raised this with two regulatory affairs specialists, both of whom regarded it as a known and unglamorous gap in how trial conduct is translated into prescribing information.3
The practice that has emerged across this class, without ever being formally codified, runs roughly as follows. Start at the initiation dose. Escalate when the current dose is comfortable enough that a person is eating normally, keeping fluids down, and not organising their week around symptoms. Do not escalate in the week of a symptom flare. If a rung is intolerable, step back to the previous one and try again later. If a rung is producing an adequate result, consider stopping there.
Every clinician we spoke to described something within a small variation of that, and none of them could point to a trial of it. It is a reasonable synthesis of the pharmacokinetics, the tachyphylaxis data and a great deal of accumulated observation.
The Journal has a specific position on this. The absence of randomised support for tolerability-led escalation is a real gap and should be closed, but its absence is not a reason to prefer the printed calendar, which has no randomised support either. Between two unrandomised schedules, the one that responds to information about the individual is the better bet. We say that as an editorial judgement rather than as a report of evidence.
In this market, gaps are usually structural rather than personal. Shortage listings, restrictions on compounded supply, price movements, customs interdiction and vendors ceasing to trade all produce interruptions that arrive without notice and end without warning. We have documented gaps of one to eleven weeks arising purely from supply, in people who missed no dose voluntarily.
The practical consequence is that anyone dependent on a single source is also dependent on that source for the continuity of their titration. Several people described re-escalating three times in a year for reasons that had nothing to do with their tolerance of the drug.
There is a second-order effect worth naming. Resuming with material from a different supplier compounds the uncertainty: the person is re-escalating and simultaneously changing the actual content of the vial. Reports from Janoshik, Medutest, PeptideMeter and VendorInvestigate consistently show that nominal strength and measured peptide content are not the same quantity, and a supplier change during a re-titration makes any symptom change uninterpretable. Change one variable at a time is a laboratory principle, and it applies here.
If there is a single practical conclusion here it is that the ladder is a default and the person is the variable. The trials that produced the figures everyone quotes were run on populations permitted to hold, delay and step back, and reading their results as an endorsement of a rigid calendar inverts what actually happened. We will keep making that point until the labels catch up with the protocols.
The evidence on stopping is better than the evidence on almost anything else in this field, because somebody deliberately randomised it.
The regain trajectories, arm by arm, with the estimands named.
The evidence base is thin and the document says so, which is to its credit.
Two withdrawal-design trials tell us what happens when treatment stops. Neither tells us what the lowest effective maintenance dose is.
Chromatographic software does not integrate every fluctuation in the baseline. It applies a threshold, and the threshold changes the reported purity by amounts that matter…
A purity figure is silent on peptide content, on water, on counter-ion, on sterility, on endotoxin and on stability. Each of those silences has a price attached.