Why does one semaglutide trial show 609 and another 30?

Aug 27 6940 views 50 posts

I keep a spreadsheet of semaglutide trial listings and I'm confused by the counts. NCT07401992 says recruiting with 62 enrolled. NCT06897475 says recruiting with 200. NCT06989203 says recruiting with 140. Then NCT02079870 is completed with 30, and NCT07011667 is active, not recruiting with 609. Why is the completed one so small while a recruiting one is 609? Am I misreading estimated vs actual enrollment? Not asking for sourcing, just trying to understand the registry fields.

👍 1

It's not just messy—enrollment can mean actual or estimated. The 30-person completed study is probably actual.

Then what about the 609 active-not-recruiting one? Is that actual or still a target?

I'd bet actual. Once it's not recruiting, they usually lock the final number.

Slow sites maybe. Some recruiting pages show current enrollment, not target.

I read the 200 one as target, not actual. ClinicalTrials is inconsistent that way.

Yeah, the 140 one might be target too. Same status doesn't mean same reporting.

👍 1

I work with registry data and the key is that enrollment can mean estimated, actual, or both depending on the field. The 609 active-not-recruiting entry likely has a final actual enrollment, while the recruiting ones may still show anticipated numbers. That doesn't tell you about dose, population, or whether it's an extension. If you're tracking semaglutide, build columns for status, phase, enrollment type, and primary completion. Otherwise the counts will look contradictory when they aren't.

Not sure phase alone explains the 30. A small completed study can still be late phase.

This is why I stopped comparing. I only looked at top-line counts and wondered why 609 dwarfed everything.

👍 1 ❤️ 1

I used to think phase told the whole story. Then I started jotting actual versus target in the margin. My little notebook now has three columns and a lot more humility.

That columns advice is gold. I added status and enrollment type to my spreadsheet last night, and suddenly the 609 and the 30 stopped fighting each other. Still confusing, but quieter.

Target, actual, or both. That's the whole ballgame.

Slow sites could explain some drift, but not a 609-to-30 gap by itself. I've seen sites post stale numbers for months while the registry sleeps. I just flag dates now.

I once spent a whole afternoon comparing five semaglutide entries. My tea went cold. The numbers never matched. Later I realized I was mixing final, estimated, and who-knows. Lesson learned.

Same status, different fields. Rookie mistake. I did that once.

The extension angle matters too. If one listing is a rollover or add-on, its count can look tiny next to a main trial. I now check parent links before side-eyeing the math.

My new rule: never compare raw enrollment numbers across semaglutide trials without a snack and a footnote. I write down phase, status, enrollment type, and primary completion. Then I close the tabs.

The 30 vs 609 thing tripped me up too until I started adding start year and phase columns to my own nerdy spreadsheet. That 30 is probably an old completed early-phase or pilot study, where they only need enough people to see whether something is worth scaling. The 609 on an active-not-recruiting record is more likely a later-phase or extension study with many sites, and the number may be estimated enrollment, not final actual. Recruiting records usually show a target, so 62/200/140 are guesses until the study closes. Registry entries also lag: completed studies sometimes never update actual enrollment, while big ones get updated by sponsors. So I stopped comparing raw counts and started comparing like-for-like: phase, start year, enrollment type, and primary completion. Weirdly, a tiny 30-person study can have more useful waist and body-comp tracking than a huge one. Does any of your rows have actual enrollment filled in, or are they all estimates?

I filter out recruiting rows entirely; active-not-recruiting still isn't results.

👍 1 ❤️ 1

Bigger n often means broader inclusion, not better body-comp data.

I made the same mistake with my own log: I compared my week 1 weight to my week 12 waist and panicked, then realized I was mixing units and timelines. Now I log the same Sunday morning, same shorts, and a waist tape. For trial records, I'd add a 'last updated' column and a 'results posted?' column. That instantly explains why old small studies look stale and big active ones look inflated.

Trial n is planned enrollment, not bodies analyzed; I learned that when a 40-person sleep study only reported 31.

👍 1 ❤️ 1

Check the actual vs estimated columns—recruiting rows default to estimated, completed ones show actual, sometimes post-dropout.

Small completed ones are often early-phase or imaging sub-studies; the 609 might be a registry-style cohort with no drug arm.

👍 2

My coed softball league gets the same treatment — an 11-tab workbook where "yes" in the RSVP column and bodies actually on the field never match up. On the 30-person completed study, poke at the Why Stopped and Locations fields: if it's one academic site, they can wrap early once they've got the small mechanistic sample they came for. The 609 could be multi-site, with several sub-studies rolled into one count. Does NCT07011667 show more than one arm or location? That's where I'd look first.

👍 2

That spreadsheet is next-level. One thing I’d check is whether enrollment means participants or observations. I got burned by that on a migraine study: n=45 was actually 15 people × 3 crossover periods. So a completed “30” could be 30 people (or 90 datapoints) depending on design, while a recruiting “609” might be planned across multiple sites or arms. Does NCT02079870 list a crossover or multiple periods in the design tab? I track my own energy/sleep in a notes app, and it’s oddly similar—same number can mean totally different things depending on the unit.

👍 5 ❤️ 1

I do my spreadsheet fiddling in airport lounges, so I feel this. One thing that tripped me up: not every NCT is an interventional drug trial. Some are observational, registries, or expanded-access, and their “enrollment” means total people observed/exposed—not randomized participants. Also one protocol can show up under multiple NCTs (parent + follow-up pieces), so counts look bonkers. Are all four of those listed as interventional? Filtering by study type plus whether there’s an arm/group column made my counts line up. If one is observational, that could explain the 609 vs 30 weirdness without anyone lying.

👍 5

The 609 might be an observational/registry study where “enrollment” counts anyone who ever touched the database, while the 30 is an interventional trial with strict follow-up. I see the same in my habit tracker: total check-ins vs fully completed weeks. If I count any log, my n looks huge; if I count perfect weeks, it’s tiny. Which number are you actually trying to compare? I’d sort by study design first. Also, tiny win—I finally color-coded my tracker so real streaks stand out.

👍 4

I’d add StartYear and Sponsor to your columns. My own trendline got less stupid when I split by year: older completed semaglutide listings cluster small because they were proof-of-concept/diabetes-era studies, while newer recruiting ones are bigger obesity/outcomes-era protocols. Also, are you deduping by sponsor + parent study? I had two NCTs that were clearly sister sub-studies, and counting both inflated my median n by ~18%. Curious what your scatter looks like if you plot enrollment vs start date instead of NCT order.

👍 7 ❤️ 1

my late-night spreadsheet habit is soda swaps, not trials, but the thing that broke my brain was “enrollment” having actual vs estimatd plus whole-study vs per-site. a completed 30 could be a tiny phase 1 or one-site sub-study, while the recruiting ones might be multi-site estimates that grow later. i started adding phase and number-of-locations columns

👍 5 ❤️ 1

New here, week one, so I’m still learning the forum ropes—but I’ve kept a little trial-notes doc for my own sanity. The 30 vs 609 thing usually sends me straight to the phase and trial start year: tiny completed ones are often early-phase or small mechanistic/PK studies, while 609 active-not-recruiting looks more like a bigger later-phase study. Also, recruiting listings often show a target number, not final enrollment. Are you sorting by phase? That’s the column that helped me stop panicking at every big number.

👍 1

Fellow chaos tracker here. Are you pulling the “actual” vs “estimated” enrollment column? Recruiting numbers are often jst site estimates, while that active-not-recruiting 609 is probably actual or rollover from an earlier phase. The completed 30 could be a small pilot or sub-study, not the same beast. I log my snack drawer the same way: planned granola bars vs actual remaining, and my kids are very unreliable sites. Also check “last update posted” — a stale recruiting page can show a tiny number that’s since been revised.

👍 4 ❤️ 1

My weird tracking trick: I add columns for phase, study design (crossover vs. parallel), and number of sites. That trio usually explains the spread without me guessing. A 30-person completed study smells like an early-phase or pilot sub-study, while big active numbers often mean a registry or extension. Do you log start year, too? Older completed trials just had smaller norms. Also, how do you keep the spreadsheet from becoming a second job? I set a 10-minute timer and only update it on coffee breaks.

👍 3 ❤️ 2

I’d add two columns: “Phase” and “Enrollment Type (Actual vs Estimated).” A completed 30-person study often smells like an early-phase PK/PD or small sub-study, while a 609-person active-not-recruiting one is likely a later-phase outcomes trial where the number is a target, not necessarily randomized. Also check “Primary Completion Date”—a tiny older study and a big newer one aren’t directly comparable even if the intervention name matches. What

👍 2

Spreadsheet goblin here too, and mine’s for lifts and macros. I’d add phase, sponsor type, site count, and primary endpoint columns. A 30-person completed study is often a PK/mechanistic or early tolerability trial at one site, while a 609 active study is probably a bigger Phase 3 outcomes thing. Enrollment numbers alone won’t tell you that. Also check the “phase” field for each—that usually explains the gap. What do yours say?

👍 4

Counts still won't line up unless you add indication and primary endpoint. I learned that tracking my own stuff: a 30-person completed trial is usually a dense mechanistic study—serial labs, imaging, food-intake diaries—so retention is brutal. 609 active-not-recruiting sounds like a multicenter outcomes/event trial, where n is driven by event rate, not recruitment ease. Also check whether enrollment means randomized, treated, or completers; those can differ by 20–30%. Are all those NCTs even in the same population? If not, stop comparing raw n. My spreadsheet has separate columns for randomized vs completed, plus indication; it killed most of my fake trends.

👍 9 ❤️ 3

Add a “number of sites” and “condition” column. A completed 30 is often one site, one narrow question — mechanistic/metabolic. A 609 active-not-recruiting is often a multi-site outcomes/registry thing pooling rollover participants. Same drug, different unit of analysis. If NCT07011667 spans 80 sites and NCT02079870 spans 1, the counts aren’t contradictory, just different experiments. What’s the condition column say for those two?

👍 4

Same spreadsheet brain, but for macros. I finally split my columns into “estimated enrollment” vs “actual enrollment” and “last registry update.” NCT02079870 has an old 2014-era prefix, so it’s probably a tiny sub-study that posted actuals. NCT07011667 active-not-recruiting likely already hit its 609 goal, so that number is closer to real. Recruiting counts are often just targets. Do you track whether results are posted and how stale the listing is? That changed my whole pivot table.

👍 4 ❤️ 1

Are you pulling the Enrollment field from the ClinicalTrials.gov API? It has enrollmentType: ESTIMATED vs ACTUAL. Recruiting records are almost always estimated; completed usually actual. That alone explains why 62/200/140 can look like real people but aren’t. Also check Start Date: NCT02079870 looks like an older early-phase study, so 30 actual could be a small pilot cohort, not a failed big trial. For the 609 active-not-recruiting one, check whether it’s an extension or rollover study—those often report total exposed across parent trials. Does your spreadsheet have enrollmentType and start year? That’s the column I’d add next.

👍 2

Fellow spreadsheet nerd—mine tracks cafe prep, not trials. The column that saved me was Enrollment Type: Actual vs Estimated. A listing can say recruiting with an estimated 200, then show actual 62 later, and the completed 30 might be actual final enrollment while 609 is still estimated. Also check Last Update Posted; registries lag and get edited. Does your sheet pull the enrollment field directly, or are you copying the headline number? If it’s the headline, that’s why 609 and 30 look like different universes. Running a cafe means I over-count every muffin I “test,” so I feel this.

👍 3

I’d separate enrollment type before anything else. On ClinicalTrials.gov, recruiting records show estimated enrollment; completed records show actual. So 62/200/140 are targets, while 30 and 609 are realized numbers. A completed n=30 is usually a Phase 1 PK/food-effect or small mechanistic study, not a mini efficacy trial. Also check if it’s event-driven: big “active, not recruiting” counts often mean an outcomes trial waiting on adjudicated events, not still enrolling bodies. Are you pulling the API or scraping the web? The API field names differ.

👍 5 ❤️ 1

We hashed this out in the spring spreadsheet thread, and the thing that finally made my sheet sane was adding phase and intervention-model columns, not just enrollment. A completed n=30 is often an early-phase or mechanistic pilot, while a 609 active-not-recruiting study is usually a bigger outcomes/extension that started years ago. Recruiting estimates also get revised upward after protocol amendments, so 62 can look tiny next to 200 even if both are early. Are you separating by phase and sponsor yet? That explained most of my weird mismatches, especially between older and newer NCT numbers.

👍 3

café owner here, and my inventory spreadsheet is the same kind of chaos. The column I’d add is “phase” + “primary purpose.” A completed n=30 is often a Phase 1 or food-effect study—small by design because they’re testing how a meal or timing affects the med, not weight outcomes. The 609 active-not-recruiting one might be a big Phase 3 where that number is planned across many countries. Does NCT02079870 list Phase 1? If so, the tiny completed count isn’t weird at all.

👍 5 ❤️ 1

I’d add a phase column and a start-year column. A completed 30-person trial is often an old early-phase or sub-study, while 609 active-not-recruiting smells like a big Phase 3 or extension. I also switch from “Estimated” to “Actual Enrollment” once results post — those numbers can drift. Does your sheet track phase, or is it just NCT + status + count?

👍 8 ❤️ 3

I know the thread’s leaning on enrollment type, but my sheet says phase/sponsor explains more of the 30 vs 609 gap. Early-phase trials are tiny by design; later ones are hundreds. Also, ClinicalTrials.gov enrollment is often anticipated, not final actual. I keep estimated vs actual as separate columns, and my completed-trial trendline usually lands 10–20% under target. What do your phase and sponsor columns show for those IDs? That’d tell you pretty fast whether you’re comparing apples to oranges.

👍 5 ❤️ 2

yo, do you have a phase column? i used to freak out comparing n’s until i added phase + condition. a completed 30 smells like an early-phase/pilot or single-site thing, while a 609 active trial is usually a bigger multi-site outcomes-ish study. also peek at the “why stopped” field on completed ones — sometimes 30 is just what finished before termination, not the original target. what phase is NCT02079870? i track my soda swaps same way: tiny n=1 trials first, then scale if it survives the week lol.

👍 5 ❤️ 2

add a phase column and a start-year delta. NCT02079870 is ancient early-phase; 30 is a pilot batch. NCT07011667 is new and probably phase 3 or observational, so 609 is a factory run.

👍 2

Check if those recruiting numbers are site-level estimates. Site pages post local targets. Registry posts whole-trial enrollment. I got burned by that on a different med. A 30-person completed trial is likely phase 1. The 609 active one is probably later phase or pooled. Also look at “Last Update Posted.” Registry data goes stale. Which NCT is phase 1 vs 3? That explains most of it.

👍 7 ❤️ 2

Also check “actual” vs “estimated” enrollment and how many sites are listed. That 30-person completed one smells like a single-center pilot/feasibility study, while a 609 active-not-recruiting count can be a multi-site extension or rollover from earlier cohorts. I run a café, and my POS does the same thing: transaction count looks huge until I filter for unique customers. If you add columns for sites, actual vs estimated, and rollover vs new enrollees, the spreadsheet stops fighting itself. Are those 609 mostly new people, or carried over?

👍 3