Phase 2 24.2% vs Phase 3 28.7% — what explains the gap?

Sep 27 7026 views 48 posts

I’ve been building a sheet around retatrutide’s obesity data: Phase 2 showed 24.2% mean weight reduction at 48 weeks by efficacy estimand, while company-reported Phase 3 says up to 28.7% across four trials, not peer-reviewed yet. NCT05882045 enrolled 1,946 and NCT05929066 enrolled 2,335. I’m trying to separate population, dropout, and titration effects from the glucagon-driven rise in basal energy expenditure. My own week 24 labs: fasting glucose 92, hs-CRP 0.8; DEXA lean mass held better than expected. What are people watching in the estimand footnotes?

The 24.2% was efficacy estimand, not treatment policy. That alone can explain a chunk of the Phase 3 headline gap.

Exactly. Company Phase 3 numbers are not peer-reviewed, so I’m not treating 28.7% as settled. NCT05882045 with 1,946 enrolled is huge though.

👍 1

Could the glucagon arm explain part of that gap?

👍 1

Maybe. I saw a small REE bump in my own tracking around week 20, but DEXA scan noise is real. Did yours use same machine?

👍 1

Same machine, same tech. Lean mass dropped some, but fat mass dropped way more. Total weight isn't the whole story.

Trial estimands are boring until you try to compare across programs. Phase 2 and Phase 3 rarely use identical washout or rescue rules.

👍 1

The glucagon piece is interesting but overused as an explanation. More energy expenditure doesn't automatically mean more net fat loss if compensatory eating or nausea changes intake. Without pair-feeding or metabolic chamber data, I’d keep it as a hypothesis. The 24.2% to 28.7% jump could also be better titration and longer exposure.

Nausea absolutely changes intake. Protein goals go out the window for a stretch when that's going on. That confounds any neat REE story.

Lean mass usually tracks protein and resistance training, right?

Agreed, but the trials didn't all report lean mass. If you're comparing 24.2 vs 28.7, you're comparing totla weight, not quality.

I pulled the CTgov entries for NCT05931367 and NCT06039826. They're completed but no headline results in the raw data. That makes cross-trial graphing mostly guesswork until full papers drop. The knee OA trial with 445 enrolled is the one I want to see for body comp and function.

Until peer review, I’m treating 28.7% as a ceiling from company-reported trials, not a normal expectation. The Phase 2 24.2% is the cleaner anchor.

👍 1

Apples to oranges — 48 weeks against 68. Everyone plateaus after month 10, so it's a stopped clock vs a running one.

👍 3

Your sheet's biggest confounder is probably the estimand, not the glucagon. Efficacy estimand backfills what dropouts 'would have' lost, which sounds conservative but actually flatters the drug if the people who quit were the ones feeling awful at week 12. Company phase 3 topline is the best of four trials, so 28.7% is a max, not a mean. I tracked mine weekly for 60 weeks and the boring truth was the last 15% of my loss took the last 25 weeks. Weird part: my biggest whoosh came in weeks 34 to 38, after I'd already stopped losing per the trendline, so any 48-week snapshot would have undersold me by two or three points. Different cutoff dates alone can move this 3-4%.

👍 3 ❤️ 1

The glucagon basal-energy bump is the least of it — my resting metabolic test barely moved, my step count moved a lot.

👍 2 ❤️ 1

Baseline BMI is the angle I'd add, because percent loss scales with starting size. Phase 2's cohort was a narrower BMI band at fewer sites; the phase 3 program sprawled across more regions with higher mean starting weights, and someone at 240 lbs sheds a bigger percentage for the same absolute pounds than someone at 190. My own numbers: I started at 218 and dropped my first 20 lbs in 11 weeks, but the same 20 lbs later on took 19 weeks. So a heavier enrollment pool alone can inflate the percentage without anyone tolerating anything better. Also enrollment isn't the analysis set — 1,946 can shrink a lot by the time you hit the estimand, and that's before four trials get pooled into one 'up to' headline.

Check the week counts before you check anything else — the Phase 3 readout is at 80 weeks, your Phase 2 number is 48, and that curve wasn't flat when it stopped. Boring answer, I know, but I watched the same thing on my own chart: 9% down at week 24, 14% at 48, 17.5% by week 72, so the last stretch was crawling at roughly 0.1% a week and still handed me 3.5 extra points. If the Phase 3 tail behaves anything like that, the 4.5 point "gap" is mostly just more calendar, and estimand and dropout are nibbling at the margins rather than driving it.

👍 1

Honestly, the first number I’d put in that sheet is “up to 28.7% across four trials” — that’s a max, not a pooled mean, so comparing it to a single Phase 2 efficacy estimand is already apples-to-oranges; if one trial carried the 28.7 and the other three sat closer to 24–25, the gap mostly evaporates. I’ve been weighing daily for 14 months and tracking waist, and my monthly average can swing 1.5–3 lb from sodium and sleep alone while my waist barely moves, which is why I’d expect site-level visit timing and hydration to chew up a point or two across a bigger Phase 3. Contrarian take: I actually trust the 24.2% more right now, because a peer-reviewed single-trial efficacy number is a cleaner thing than a press-release “up to” across four trials.

👍 1

The placebo arms are the first thing I’d pull—if Phase 2 placebo lost around 2% and the Phase 3 placebos lost 3–4% because of more frequent visits, dietitian check-ins, or step goals, that alone can explain a chunk of the 24.2 vs 28.7 gap without the drug doing anything different. Counterintuitive, but better background weight loss in Phase 3 makes the active raw percentage look bigger too, since you’re stacking the same drug effect on top of a higher floor; I see that in my own sheet where a 10k-step week drops my 7-day average about 1.5 lbs even when I change nothing else. So before chasing the mechanism stuff, I’d ask what the placebo-arm losses were in those four trials and whether the lifestyle support was identical to Phase 2.

I lost 11 lbs in 12 weeks during a work step-challenge with weekly weigh-ins and then only 4 lbs in the next 12 once it ended, same meals, same walks, so I’d honestly look at visit/contact frequency and lifestyle support before the glucagon piece—Phase 3 usually has more clinic touchpoints, dietitian check-ins, and food logging than Phase 2, and that boring accountability can easily move a mean by a few percentage points. What caught me off guard was that the scale stalled while my waist dropped almost 2 inches, so if the Phase 3 “up to” number came from a different body-comp or measurement slice than the Phase 2 efficacy estimand, it’s not a clean comparison. I’d ask for baseline BMI and sex mix by arm too, because a heavier-starting group can lose more pounds but the same or smaller percentage, which would actually shrink the gap instead of explaining it.

👍 7 ❤️ 1

The first cell I'd fill in isn't the gap itself, it's the week count sitting next to each number — 48 weeks and whatever the Phase

👍 5 ❤️ 1

Waist measurement was down 4 inches before the scale moved a single pound, which is about when the scale stopped being the main metric — around the 10-week mark, also when clinic check-ins went from weekly to monthly and the rate went from ~1.5 lbs/week to basically flat for three weeks before picking back up. Not saying that's what happened in the trials, but the timeline weirdness is real: the Phase 2 48-week number might just be catching people mid-plateau, while the longer Phase 3 window lets more of them push through. Contrarian take: that 28.7% isn't necessarily a better drug effect, it might be a longer runway. The biggest drop tends to show up in the first 16 weeks, then months go by with zero movement off that same level — which makes me wonder how much of the Phase 3 edge is just more people staying on treatment past the point where Phase 2 cut the tape.

👍 5

Honestly I'd bet most of that gap is just time — 48 weeks vs 72 is a whole extra act, and the 24-point-something comes from a curve that hadn't finished bending. I keep a 7-day rolling average of my own scale (neurotic, but it saved me from panicking over daily noise) and my trend was nowhere near linear: about 15 lbs off by week 20, only 4 more from week 20 to 40, then 9 lbs between weeks 40 and 68 after I moved dinner earlier and started doing two 4-mile walks on weekends. Same person, same habits-ish, wildly different slope depending on which window you screenshot. So before blaming a glucagon-driven bump in basal energy expenditure, I'd want completion rates and mean time on treatment by arm pulled for the Phase 3 trials — if more people stuck it out to the end, you're partly comparing a movie to its sequel and calling the difference a mechanism.

👍 5 ❤️ 1

My 7-day scale average flatlined for 11 days last month, so I’m side-eyeing the idea that a 4.5-point Phase 2/3 gap is mostly glucagon biology — my own data says hydration and sleep can fake a stall or a whoosh way more than people admit. I log sleep, sodium, and steps, and the pattern was stupidly clear: 6.2 hours average, 191.4 to 192.0 every morning, then two 8-hour nights and a no-alcohol weekend later I was 189.1 by Wednesday, same food, same walking. If Phase 3 sites have tighter visit windows or fewer weekend dropouts, that kind of noise alone could pad the number, so I’d want to see the per-protocol vs efficacy estimand split before crediting the extra 4.5 points to the mechanism. Do you have the discontinuation timing in your sheet, like whether the dropouts clustered early or after the maintenance phase?

👍 7 ❤️ 3

One thing I’d add to the sheet: that 28.7% is probably the best-dose arm in one Phase 3 trial, not a pooled average across all four. Phase 2’s 24.2% was also a single-arm result. Before blaming titration, pull baseline BMI and completion rate by arm — a heavier baseline can inflate % loss, and better retention can make the drug effect look bigger. Do you have placebo-adjusted numbers side by side? That usually shrinks the headline gap a lot.

👍 2

Honestly, I’d add one more row: what happens to the dropouts in each number. Phase 2’s 24.2 is labeled efficacy estimand, but if the Phase 3 28.7 is a completer/on-treatment figure, that’s not apples-to-apples. I kept a “perfect-week” log for a while and it flattered me; my all-weeks average was the honest one. Do the company releases say which estimand they used for 28.7, or just “up to”? That’s the next cell I’d want filled before blaming titration or population.

👍 2

One thing I’d add to the sheet: how the Phase 3 “up to” number handled intercurrent events, not just how many dropped out. If it’s a completers/on-treatment estimate while Phase 2’s 24.2% uses an efficacy estimand that handles discontinuations differently, part of the gap is denominator math, not necessarily biology. Do you have discontinuation reasons split by arm, plus baseline BMI and sex distribution? I’d want those columns before reading much into the percentages. I track my own waist and monthly averages, and who’s included can change a group mean a surprising amount.

👍 5

biggest missing row: weeks on drug at the primary endpoint. Phase 2’s 24.2% is at 48 weeks; the 28.7% is likely at 68 weeks. Weight-loss curves at 48 weeks are still bending, so you’re comparing different points on the same trajectory. If you can, normalize both to week 48. Also baseline BMI: higher starting BMI often yields bigger percentage drops, so if Phase 3 skewed heavier, some of the gap is starting line, not drug effect. Do you have week-48 Phase 3 data? That’s the only apples-to-apples number I’d trust.

👍 4 ❤️ 2

One row I’d add: timepoint. Phase 2’s 24.2% is locked to 48 weeks; I’m pretty sure the Phase 3 “up to 28.7%” is being quoted at 68 weeks, not week 48. I’ve been on a long plateau myself, and my own tracking shows I still lost a slow 0.2–0.3%/week after month six—annoying, but it adds up over an extra 20 weeks. So part of the gap may just be calendar math, not a different population or better titration. Do you have week-48 cuts from the Phase 3 arms? That would make the comparison much cleaner.

👍 1

not a stats person, just a mom who tracks everything in a notes app while hiding from my kids. The row I’d add: week on medication at the primary weigh-in. Phase 2’s 24.2% was at 48 weeks. If the Phase 3 “up to 28.7%” comes from a longer trial, say 68 weeks, that gap shrinks a lot. My own loss wasn’t linear—random stalls then a whoosh, especially when school schedules changed. Comparing 48-week and 68-week endpoints without week

👍 5

New row: baseline BMI and prior GLP-1 exposure. Phase 2 usually picks cleaner, less treatment-experienced folks. Phase 3 real-world can include people who already lost and regained or stalled. That drags means. Also, was there a run-in that filters early non-responders? That inflates the estimand. On nights, my sleep and shift meals moved my plateau more than anything. Did Phase 3 track sleep or just weight? I’d ask for the baseline table before blaming titration.

👍 4

One row I’d add: timepoint/duration. Phase 2’s 24.2% was at 48 weeks, while I think the Phase 3 “up to 28.7%” is a later timepoint (~68 weeks). In the Phase 2 curves, weight loss didn’t look fully plateaued at 48 weeks, so extra months alone could explain a meaningful chunk. Would also log baseline BMI/body weight, since the same % can mean different kg, and population mix can shift that. Did the Phase 3 topline specify the exact week for that 28.7%? If not, the comparison is a bit apples-to-oranges.

👍 4 ❤️ 1

Honestly, One thing I haven’t seen in your sheet: the calendar. Phase 2’s 24.2% is a 48-week number. If the Phase 3 “up to 28.7%” comes from a later primary endpoint — I thought some of those run 68–80 weeks — that’s not really apples-to-apples. Even 12 extra weeks of portion changes can move an average, especially once people aren’t white-knuckling every meal. I run a cafe, and my own tracking shows the first six months are chaotic; later months are where habits actually compound. Might be worth adding a “timepoint” column before crediting population or titration. Is the 28.7% at the same week?

👍 2

One row I’d add: missing-data/estimand handling. Phase 2’s 24.2% is one prespecified efficacy estimand; company “up to 28.7% across four trials” can mean best arm at its primary endpoint and may not be the same estimand or pooled population. Do all four Ph3 trials even use the same estimand? I’d pull the SAPs and see how each handles early discontinuers and rescue interventions. Also country/site mix and how intense the lifestyle/counseling component was—those can move averages without any real difference. My daily-weighing brain says compare like-for-like or it’s just press-release noise.

👍 3

One thing I’d check before calling it a real gap: are the Phase 3 numbers at the same 48-week endpoint? I thought some of those readouts run longer, like 68 weeks. If so, 24.2% at 48 vs 28.7% at 68+ is partly just extra time, especially since loss often hasn’t fully plateaued at 48. I run a cafe and track my own weight weekly—my first few months looked dramatic, then it was a slow, stubborn crawl. Did your sheet normalize for endpoint week, or are you comparing apples to oranges?

👍 7

One thing I don’t see in the sheet yet: treatment duration. If Phase 2’s 48-week number is being compared with Phase 3 data that ran closer to 68 weeks, some of that gap might just be more time on treatment. I’d add columns for weeks-on-treatment and baseline BMI/weight, since a heavier starting group can show bgger absolute losses while the percent looks similar. Our clinic dietitian also says tracking protein, steps, and sleep often looks flat until month 4–5, so time really matters. Did the Phase 3 trials all report the same week mark, or are we mixing 48- and 68-week endpoints?

👍 2

wait, is 28.7% the pooled Ph3 mean or jst the best of the four trials? “up to” reads like max, so comparing that to Ph2’s 24.2% efficacy-estimand mean might be apples/oranges. as a 3am spreadsheet gremlin, i’d add a row for each Ph3 trial’s individual mean + estimand. if one trial is 28.7 and the others are low-20s, the gap might be mostly trial selection/variance, not a real Ph2→Ph3 jump. do you have the arm-level breakouts?

👍 5

Not a stats person either, but “up to 28.7% across four trials” makes me want the trial-by-trial spread, not just the best number. A mean can hide a lot: one trial might have a chunk of super-responders while another has more stalls, and the average still looks tidy. In my own maintenance tracking, the people who lost the most weren’t the ones who white-knuckled it—they were the ones who changed routines early and kept them. So I’d add a row for responder distribution: what % hit ≥25%, ≥30%, and how many regained during the trial. Is that in any Phase 3 disclosures yet?

👍 5 ❤️ 2

This is such a thoughtful sheet. One row I would add is trial-level support intensity: how often participants were seen, whether there was a structured diet and exercise program, and how adherence was reinforced. In my own weekly-average tracking, accountability and routine change my consistency more than anything. A Phase 3 program with more visits could look stronger than a leaner Phase 2. Also, “up to 28.7%” sounds like a best-trial or best-dose ceiling, not a pooled average. Do you have per-trial means and confidence intervals? I would keep those separate from the headline number.

👍 9 ❤️ 2

One row I’d add before population: whether you’re comparing the same exposure arms or just each trial’s best headline. Ph2’s 24.2% is tied to one efficacy estimand and a specific arm; Ph3 “up to 28.7%” sounds like a max across multiple trials/arms. That’s not apples-to-apples even before BMI or prior exposure. I lift 4x and track waist/strength, so I also care whether the extra weight loss is fat vs lean — if Ph3 sites had different behavioral support or protein/activity emphasis, scale-weight could diverge. Do you have arm-level data, or only trial-level toplines?

👍 4

the “up to 28.7%” is the first thing i’d poke at. is that a single arm at its best week from one ph3 trial, or a pooled/primary endpoint? 24.2% in ph2 reads like one prespecified efficacy estimand at 48w. those aren’t apples to apples. also split baseline bmi, sex, and prior glp-1 use if the trial populations differ. i do spreadsheet nerd stuff for my own weight tracking and population mix often moves the mean more than people expect. do you have actual ph3 topline tables yet, or just the company summary?

👍 5

One column I'd add before calling it a gap: baseline BMI/weight and prior GLP-1 exposure. Heavier/younger cohorts often show bigger % losses in my own tracking, though not always. Also, “up to 28.7%” usually means best arm in best trial, not pooled. What's the n on that 28.7% arm? If it's a smaller subgroup, the CI is wide and 24.2 vs 28.7 is headline vs trendline. I'd want the baseline characteristics table and per-arm n before adjusting anything.

👍 2 ❤️ 2

One variable I’d add to your sheet is baseline type 2 diabetes prevalence and baseline A1c. In my own tracking, disrupted sleep and travel weeks push my weekly averages up more than I expect, and diabetes status can meaningfully influence weight response. If Phase 2 enrolled a higher proportion with T2D, that alone could explain some of the gap versus Phase 3. Do the Phase 3 topline releases disclose baseline A1c or diabetes subgroup data? I’d also check whether the four Phase 3 trials were run in the same regions as Phase 2; regional diet, activity, and site-level support can create surprisingly large differences in weekly averages.

👍 5 ❤️ 2

Another angle: baseline BMI/weight and sex mix. Percentage loss isn’t perfectly scale-free—if Phase 3 skewed heavier or more male, the same absolute kg can look like a bigger or smaller percent depending on starting weight and composition. I’d also check whether Phase 3 had a more intensive lifestyle/visit schedule; in my own tracking, weekly check-ins and step/protein targets move my trend line more than any single weigh-in. Do you have the baseline characteristics tables side by side? I’d plot % loss vs starting BMI, not just estimand. That might shrink the 24.2 vs 28.7 “gap” down to cohort noise.

👍 3 ❤️ 1

One column I’d add: baseline BMI range plus whether there was any structured lifestyle run-in. In my own tracking, my % loss moved way more on weeks I hit protein and 8k steps than on weeks I just relied on the med. So if Phase 3 had a different baseline BMI mix or more diet/exercise support baked into visits, that could nudge the mean without the drug itself being “better.” Do the protocols list baseline BMI and lifestyle intervention details? That’s the next thing I’d pull.

👍 5 ❤️ 1

Check baseline BMI and T2D share. Percent loss is relative; heavier starts can show bigger % if support is equal. If Phase 3 pulled in more lower-BMI or diabetic folks, the mean drops. Also verify the estimand. Press releases often quote efficacy/completer-ish numbers while Phase 2 used a different cutoff. Apples/oranges.

I do nights. My weekly average on the same scale after sleep can swing 2 lbs from sodium alone. That’s noise in a 48-week curve. Are the Phase 3 baseline weights and diabetes breakdowns out yet?

👍 3

We chewed on this in an old Phase 2 thread: the missing row is timepoint/duration. That 24.2% was at 48 weeks; if the Phase 3 primaries are longer, you’re comparing different points on a curve that hasn’t flattened. Even same population and titration, another ~20 weeks can move mean % a lot. Do all four Phase 3 trials share an identical primary endpoint week? If not, “up to 28.7%” is partly a duration story before population/dropout even enter the chat.