SURMOUNT-1 extension data: what happens after the trial stops
A design note rather than a result: what the comparator was, and what that permits you to conclude.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Measurement
Density is a proxy for strength and an imperfect one, particularly when soft-tissue thickness over the measurement site is changing.
The mechanistic argument for a genuine skeletal effect is simple and reasonable: bone remodels in response to mechanical loading, a lighter person loads their skeleton less, and a rapid reduction in loading produces a rapid reduction in density. That argument predicts bone loss with any successful weight-loss intervention and does not predict anything peculiar to incretins. Whether these drugs do something to bone beyond making their users lighter is a separate question and the evidence for it is currently very thin in both directions.
The closest analogue to rapid weight loss in an older, heavier population predates this drug class entirely. In a randomised trial of adults aged sixty-five and over with obesity, assigned to diet, exercise, both or a control condition for a year, the combination produced the largest improvement in physical function, and the exercise component attenuated the loss of lean mass and of bone mineral density that diet alone caused.1 Diet alone improved function too — carrying less mass helps — but by less, and at a measurable skeletal cost.
That trial is the template for how the question should be asked in this class: randomise the co-intervention, measure function as a primary endpoint, measure bone, and follow for long enough for the skeleton to respond. Its population, older and heavier and losing weight quickly, resembles a large share of current incretin users far more closely than the young resistance-trained cohorts from which most consumer advice descends.
The Journal cites it frequently for that reason and notes the obvious limitation: the weight loss achieved was roughly a tenth of body mass over a year, which is half or less of what the current agents produce. Whether the protective effect of training holds at twice the rate of loss is not established.
A secondary analysis of the Danish exercise-and-liraglutide trial is the only randomised evidence on bone in this class worth the name. It reported that exercise alone, or exercise combined with the agonist, preserved bone mineral density at clinically relevant sites, whereas the agonist alone was associated with reductions at the hip and spine relative to the exercise arms.2 The effect sizes are small in absolute terms and the trial was not designed for this endpoint.
Around that sits a larger and older literature on dietary and surgical weight loss, which is consistent: substantial weight reduction lowers bone mineral density at load-bearing sites roughly in proportion to the mass lost, with the hip and femoral neck affected more than the lumbar spine, and with bariatric surgery producing the largest changes. Bone turnover markers rise early and remain elevated for months.
Two things are missing. There is no randomised bone endpoint in any trial of the current agents, at any dose, for any duration. And there is no fracture data at all — no trial in this class has been powered for fractures, none has reported them as a pre-specified outcome, and the observational literature is confounded by the fact that weight loss changes fall risk in both directions.
Three hundred scanned participants are carrying the entire public argument about whether this drug class costs its users muscle.
On the substudy evidence baseDensitometry infers bone mineral density from the differential attenuation of two X-ray energies, using the surrounding soft tissue as the baseline against which bone is distinguished. The algorithm assumes a soft-tissue composition, and that assumption is embedded in the calibration. When the thickness and fat fraction of the tissue overlying a measurement site change substantially, part of the apparent change in bone density is an artefact of the altered baseline.
The magnitude is contested. Phantom and cadaver work suggests errors of the order of one to three per cent for large changes in overlying fat, which is the same order as the real bone changes being reported over a year of rapid weight loss. In practice this means that a hip bone mineral density reduction of two per cent in a person who has lost a fifth of their body weight cannot be cleanly separated into a bone effect and a measurement effect, and the published analyses do not attempt it.
Quantitative computed tomography and high-resolution peripheral imaging are less vulnerable, measure geometry and microarchitecture rather than areal density, and have not been used in any trial in this class. The Journal regards that as the most easily closed gap in the whole body-composition literature.
| Endpoint | Measured in a randomised trial? | Where |
|---|---|---|
| Areal BMD, hip and spine | Yes, as a secondary analysis | S-LiTE bone analysis |
| Bone turnover markers | Yes, small studies | Investigator-initiated |
| Bone geometry or microarchitecture | No | — |
| Incident fracture | No | — |
| Falls | No | — |
| Absence from this table means the Journal could not find a pre-specified randomised measurement, not that no observational data exists. Observational fracture data in weight loss is confounded in both directions. | ||
Two hypotheses compete and both are underpowered. The first is that incretins are neutral for bone beyond making their users lighter, so any density change is the ordinary consequence of reduced mechanical loading. The second is that GLP-1 receptor signalling has direct skeletal effects — receptors have been reported on osteoblast lineage cells, and GLP-1 influences the entero-osseous axis and calcitonin secretion — which could be protective, harmful, or negligible.
The evidence cited for a protective effect is an early study of weight-loss maintenance in which liraglutide treatment was associated with preserved bone mineral density relative to a diet-alone comparison, interpreted at the time as a direct skeletal benefit.3 That finding sits awkwardly beside the later secondary analysis in which the agonist arm did worse than the exercise arms, and the two are not straightforwardly reconcilable: different agents at different doses, different comparators, different durations, small samples throughout.
The Journal reports the question as open, which is unsatisfying and accurate. What would settle it is a randomised bone endpoint with imaging that is not confounded by soft-tissue change, in a population whose weight loss is matched across arms. Nothing of that description is under way.
Four things accompany every composition number in these pages. The instrument, because DXA, magnetic resonance, bioimpedance and creatine dilution are not interchangeable and the choice frequently determines the sign of the result. The sample size of the substudy rather than of the parent trial, because the parent trial size is irrelevant to the composition finding and quoting it is misleading. The definition used — total lean mass, lean soft tissue, appendicular lean mass or fat-free mass — because these differ by several kilograms in the same person. And whether the figure is a proportion of body mass or an absolute quantity.
Where a source omits any of the four, we say so rather than guessing, and where we have had to convert between definitions we show the conversion. This is more cumbersome than the alternative and it is the only way we have found to write about this subject without producing sentences that are technically true and practically misleading.
Readers who find a figure in these pages that lacks its instrument and its sample size have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
A category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.
A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.
The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.
Correspondence on this subject reaches the Journal at a higher rate than on any other, and a striking proportion of it consists of readers reporting a number from a device and asking what it means. The honest answer, in most cases, is less than they hope. We would rather say that than supply a confident interpretation the instrument cannot support.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
I have read your protein tables twice and I still cannot work out what I should eat. I appreciate that this is the honest position but it is not a useful one for a person in a supermarket.
— D. Ferreira-Lopes, Porto
It is a fair complaint about a real limitation. What we can say is that the defensible range is narrower than the disagreement suggests, that the denominator matters more than the ratio, and that a clinician or dietitian can convert a range into a number for your body in a way that a magazine cannot.
Three vendors have now sent me marketing material claiming their product preserves lean mass during GLP-1 treatment, two of them citing your publication as a source for the underlying composition figures. You may want to know that.
— A. Petrucci, Bari
We did not, and we are grateful. Quoting our reporting of a substudy alongside an unevidenced product claim is a misuse of it, and the standards desk has written to all three.
I am sixty-eight, I have lost nineteen kilograms over fourteen months, and my consultant has twice told me my lean mass is fine on the basis of a handheld bioimpedance device in the clinic corridor. Having read your piece on what that device measures, I am no longer sure what I have been reassured about.
— B. Ademola, Ilorin
Nor are we. A handheld device measures impedance across the upper body and infers the rest, and the inference is least reliable exactly where you sit: older, substantial weight change, changing hydration. That is not a criticism of your consultant’s judgement, which may be sound on other grounds, but the device is not the evidence for it.
Small correction to your table: the S-LiTE exercise prescription was two supervised group sessions and two individual sessions weekly, not two sessions in total. The distinction matters because "add some exercise" is not what was tested.
— F. Aubert, Toulouse
Correct, and that is precisely the point we were trying to make and then undermined in our own table. Amended.
You write that no trial has measured strength. There are observational cohorts with grip strength data. Why do you insist on randomised measurement?
— M. Sandhu, Amritsar
Because grip strength in an observational cohort of people who chose to take a drug, and who differ from those who did not in age, motivation and comorbidity, cannot separate the drug effect from the selection. We report those cohorts and we do not treat them as answering the question.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Where the familiar figures come from, what populations they were measured in, and how far the extrapolation reaches.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
A survey of what the meta-analyses support, with the populations named.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
A design note rather than a result: what the comparator was, and what that permits you to conclude.