What happens when the drug is taken away: three randomised answers
What was withdrawn, from whom, after how long, and what was measured afterwards.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Skeletal health
One randomised trial has combined a GLP-1 receptor agonist with supervised exercise. Its result is the single most useful piece of evidence in this area.
The most instructive trials in this area were run before anybody had heard of an incretin. In older adults with obesity randomised to diet, exercise, both or neither, the combination preserved physical function and attenuated the loss of bone and lean tissue that dieting alone produced. That population — older, heavier, losing weight fast — resembles a substantial share of the people now taking these drugs far more closely than the resistance-training cohorts from which most protein and training advice is drawn.
Dual-energy X-ray absorptiometry is the reference method in this field for practical rather than theoretical reasons: it is fast, the radiation dose is trivial, it is widely installed, and it reports regional as well as whole-body values. Its coefficient of variation for whole-body lean mass on a well-maintained clinical scanner with a consistent operator is on the order of one per cent, which sounds excellent until it is converted into kilograms. For a person with fifty-five kilograms of lean tissue, a one per cent coefficient of variation implies a least significant change — the smallest difference between two scans that can be distinguished from measurement noise with reasonable confidence — of roughly one and a half kilograms.
Appendicular lean mass, the arms-and-legs subtotal that is the closest DXA proxy for skeletal muscle, has a smaller absolute magnitude and a somewhat larger relative error, and the two effects roughly cancel. Regional values for a single limb are noisier again. None of this is a criticism of the instrument. It is the reason a body-composition report that changes by half a kilogram between visits has told the person nothing, and the reason the trial substudies report group means rather than individual trajectories.
An imaging substudy inside a large trial is sized to describe rather than to test. The enrolment is set by how many participating sites have a scanner and by what the sponsor budgeted, not by a power calculation against a composition hypothesis, and the analysis is generally pre-specified as exploratory or descriptive. The consequence is that these substudies can report a mean change with a usable confidence interval and cannot support most of the questions asked of them.
They cannot, for instance, establish whether lean-mass change differs between dose arms, because the per-arm enrolment after splitting is in the low tens. They cannot establish whether it differs by age, sex, baseline adiposity or diabetes status, because those subgroups were not enrolled to be comparable. They cannot describe the distribution of individual responses, because the per-participant least significant change is a substantial fraction of the observed mean effect. And they cannot address function at all, because nobody measured it.
Nor was the imaging repeated when the programmes were extended. The two-year semaglutide extension reported weight, waist circumference and cardiometabolic parameters at week 104 and did not repeat the composition substudy, so there is no imaging at all beyond seventy-two weeks in this class.1 Whatever the trajectory of lean mass is in year two of treatment, nobody has measured it.
None of this is a scandal; it is the ordinary economics of trial substudies. It becomes a problem only when a descriptive group mean is quoted as though it characterised what will happen to an individual, which is now the normal register of coverage on this subject.
The instrument determines the answer more than the drug does, and the trade quotes the answer without naming the instrument.
Priya Ramanathan, PharmD, Pharmacy ColumnistA Danish randomised trial remains the only controlled test of the obvious question. After an eight-week low-energy diet producing approximately thirteen kilograms of weight loss, participants were randomised for one year to supervised exercise alone, liraglutide 3.0 mg alone, both combined, or placebo.2 The combination arm achieved the largest weight reduction and, more relevantly here, the most favourable composition outcome: body fat percentage fell roughly twice as much in the combination group as in either single-intervention group, and the exercise arms preserved lean mass better than the drug-alone arm.
Three qualifications belong with that result. The exercise was supervised and substantial — two group sessions and two individual sessions weekly, with a vigorous-intensity target — which is not what most people mean by adding exercise. The agent was liraglutide at 3.0 mg daily, producing considerably less weight loss than the current agents, so whether the interaction scales to a twenty per cent reduction is unknown. And the trial began after weight had already been lost, so it is a maintenance study rather than an induction study.
With those stated, it is the best evidence in the field and it points in the direction the general advice already points.
| Endpoint | Measured in a randomised trial? | Where |
|---|---|---|
| Areal BMD, hip and spine | Yes, as a secondary analysis | S-LiTE bone analysis |
| Bone turnover markers | Yes, small studies | Investigator-initiated |
| Bone geometry or microarchitecture | No | — |
| Incident fracture | No | — |
| Falls | No | — |
| Absence from this table means the Journal could not find a pre-specified randomised measurement, not that no observational data exists. Observational fracture data in weight loss is confounded in both directions. | ||
The closest analogue to rapid weight loss in an older, heavier population predates this drug class entirely. In a randomised trial of adults aged sixty-five and over with obesity, assigned to diet, exercise, both or a control condition for a year, the combination produced the largest improvement in physical function, and the exercise component attenuated the loss of lean mass and of bone mineral density that diet alone caused.3 Diet alone improved function too — carrying less mass helps — but by less, and at a measurable skeletal cost.
That trial is the template for how the question should be asked in this class: randomise the co-intervention, measure function as a primary endpoint, measure bone, and follow for long enough for the skeleton to respond. Its population, older and heavier and losing weight quickly, resembles a large share of current incretin users far more closely than the young resistance-trained cohorts from which most consumer advice descends.
The Journal cites it frequently for that reason and notes the obvious limitation: the weight loss achieved was roughly a tenth of body mass over a year, which is half or less of what the current agents produce. Whether the protective effect of training holds at twice the rate of loss is not established.
Two claims are routinely bundled together and only one is well supported. The weaker claim is that resistance training during pharmacological weight loss builds or maintains muscle mass. In a substantial energy deficit, training generally attenuates the loss rather than preventing it, and net accrual is unusual outside of untrained beginners and the specific controlled-feeding conditions of the trials cited earlier. The stronger claim is that training preserves strength and physical function even where mass declines, which is consistently observed and is mechanistically sensible: a large part of early strength change is neural rather than structural.
The distinction has practical consequences. Somebody training hard, eating well, and watching their DXA appendicular lean mass fall by two kilograms across nine months has not failed at anything, and may be measurably stronger than at baseline. If the expectation set for them was mass preservation, they will read a normal outcome as a failure and may respond by eating more or training in ways that suit the metric rather than the goal.
The Journal reports the training recommendation and reports what it is expected to achieve, which is function first and mass second.
The word sarcopenia has migrated from clinical medicine into consumer discussion of this drug class and lost its definition in transit. In the working definitions used by the European and Asian consensus groups, sarcopenia requires low muscle strength, with low muscle quantity or quality confirming it and poor physical performance indicating severity. Strength is the entry criterion. Reduced lean mass on a scan, in the absence of measured weakness, does not meet any published definition of sarcopenia.
This matters because the borrowed term imports a prognosis. Sarcopenia in its clinical sense is associated with falls, fractures, hospitalisation and mortality, and those associations were established in older adults with measured weakness, frequently in the context of illness or immobility. Applying the label to a forty-two-year-old whose DXA appendicular lean mass has fallen by one and a half kilograms while their strength has increased is not a cautious extrapolation; it is a category error with a frightening prognosis attached.
The related term sarcopenic obesity has the same problem in a more acute form, since it requires both criteria to be met and is frequently used to mean nothing more than a low lean fraction. The Journal uses both terms only in their defined sense and asks correspondents who use them to say which criteria they mean.
Four things accompany every composition number in these pages. The instrument, because DXA, magnetic resonance, bioimpedance and creatine dilution are not interchangeable and the choice frequently determines the sign of the result. The sample size of the substudy rather than of the parent trial, because the parent trial size is irrelevant to the composition finding and quoting it is misleading. The definition used — total lean mass, lean soft tissue, appendicular lean mass or fat-free mass — because these differ by several kilograms in the same person. And whether the figure is a proportion of body mass or an absolute quantity.
Where a source omits any of the four, we say so rather than guessing, and where we have had to convert between definitions we show the conversion. This is more cumbersome than the alternative and it is the only way we have found to write about this subject without producing sentences that are technically true and practically misleading.
Readers who find a figure in these pages that lacks its instrument and its sample size have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
No head-to-head trial has compared body composition between agents in this class. Every published ranking is an artefact of the comparison.
On the muscle-sparing claimA category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.
A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.
The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.
| Programme | Agent | Method | Substudy n (approx.) | Duration |
|---|---|---|---|---|
| STEP 1 | Semaglutide 2.4 mg | DXA, whole body | 140 | 68 weeks |
| SURMOUNT-1 | Tirzepatide 5/10/15 mg | DXA, whole body | 160 | 72 weeks |
| SURPASS-3 MRI | Tirzepatide vs degludec | MRI, liver and abdominal depots | 300 | 52 weeks |
| S-LiTE (investigator-initiated) | Liraglutide 3.0 mg ± exercise | DXA, whole body and regional | 195 | 52 weeks |
| SURMOUNT-4 | Tirzepatide, withdrawal design | No imaging substudy reported | — | 88 weeks |
| Enrolment figures are approximate and refer to the imaging substudy, not the parent trial. Substudy sites were selected for scanner availability rather than for representativeness. | ||||
The Journal position on body composition is narrower than either camp would like. The drugs produce weight loss whose composition is, on the available imaging, at least as favourable as dietary weight loss and probably slightly better. The absolute lean-mass reduction accompanying a twenty per cent weight loss is nevertheless substantial, is unmeasured in functional terms, and is a reasonable thing for an older or already frail person to want managed. Both of those sentences are true and the argument has largely consisted of people insisting on one of them.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
I have read your protein tables twice and I still cannot work out what I should eat. I appreciate that this is the honest position but it is not a useful one for a person in a supermarket.
— N. Prasetyo, Surabaya
It is a fair complaint about a real limitation. What we can say is that the defensible range is narrower than the disagreement suggests, that the denominator matters more than the ratio, and that a clinician or dietitian can convert a range into a number for your body in a way that a magazine cannot.
Three vendors have now sent me marketing material claiming their product preserves lean mass during GLP-1 treatment, two of them citing your publication as a source for the underlying composition figures. You may want to know that.
— V. Bhattarai, Kathmandu
We did not, and we are grateful. Quoting our reporting of a substudy alongside an unevidenced product claim is a misuse of it, and the standards desk has written to all three.
The soft-tissue artefact point in your bone section is underplayed. In a patient losing twenty per cent of body mass the change in overlying tissue is well outside the range the calibration was validated over, and the published analyses do not report a sensitivity analysis for it. That is not a caveat, it is a gap.
— P. Sarkissian, Beirut
We accept the escalation and have strengthened the wording. The absence of any published sensitivity analysis is, as you say, the more damaging observation.
My mother is eighty-one and on a low dose for her diabetes. Her weight is down nine kilograms and she now struggles to get out of a low chair, which she did not eighteen months ago. Nobody has measured anything. I do not know whether this is the drug, the weight loss, or being eighty-one, and neither does anybody I have asked.
— J. Delahunty, Waterford
That is the situation the missing endpoint produces, and we are sorry to have no better answer. A chair-stand time takes thirty seconds to measure and would at least establish a baseline against which the next six months could be judged. It is worth asking for by name.
You write that no trial has measured strength. There are observational cohorts with grip strength data. Why do you insist on randomised measurement?
— E. Nkomo, Polokwane
Because grip strength in an observational cohort of people who chose to take a drug, and who differ from those who did not in age, motivation and comorbidity, cannot separate the drug effect from the selection. We report those cohorts and we do not treat them as answering the question.
What was withdrawn, from whom, after how long, and what was measured afterwards.
Escalation beyond the label is common in this market. Reporting that it happens is not the same as reporting that it works.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Density is a proxy for strength and an imperfect one, particularly when soft-tissue thickness over the measurement site is changing.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Three randomised withdrawal designs have tested what happens when treatment stops. Their results are consistent and they are consistently misreported.