[LDN] in Crohn’s Disease: A Prematurely Terminated Randomized Trial Showing Reduced Fatigue Without Clinical or Endoscopic Benefit 2026 van de Pol+

Andy

Senior Member (Voting rights)
Full title: Low-Dose Naltrexone in Crohn’s Disease: A Prematurely Terminated Randomized Trial Showing Reduced Fatigue Without Clinical or Endoscopic Benefit

Background​

Crohn’s disease (CD) is often refractory to standard therapies, and patients are sometimes reluctant to initiate immunosuppressive therapy, prompting interest in novel treatments such as low-dose naltrexone (LDN).

Aims​

To evaluate the efficacy of LDN for induction of remission in patients with mild-to-moderate active CD.

Methods​

We conducted a multicenter, double-blind, placebo-controlled trial in 7 hospitals in the Netherlands. Adults with active CD (defined by mucosal ulcers on endoscopy and Simple endoscopic score for CD 3–15) were randomized 1:1 to LDN 4.5 mg once daily or placebo for 12 weeks. The primary endpoint was endoscopic remission at week 12. Secondary endpoints included clinical and endoscopic response, safety and patient-reported outcome measures. The trial was stopped early due to futility.

Results​

Forty-one patients were randomized. After 12 weeks, endoscopic remission occurred in 1 (5.6%) vs. 3 (18.8%) patients for LDN and placebo (p=0.233), respectively. Clinical remission occurred in 22.2% (LDN) vs. 70.0% (placebo) (p=0.037). No serious adverse events were reported. For the FACIT-Fatigue questionnaire, the median difference after 12 weeks was 2.5 (IQR 0–7) for LDN vs. −3 (IQR −7–2) for placebo (p=0.012). Additionally, patient-reported outcomes showed no significant difference between LDN and placebo.

Conclusions​

LDN was not effective for inducing remission in CD and showed no benefit over placebo. Despite not being an effective treatment for clinical and endoscopic outcomes, it may enhance better quality of life by reducing fatigue symptoms. Future studies are needed to clarify its potential benefit on fatigue in patients with inflammatory bowel disease.

Open access
 
After 12 weeks, endoscopic remission occurred in 1 (5.6%) vs. 3 (18.8%) patients for LDN and placebo (p=0.233), respectively. Clinical remission occurred in 22.2% (LDN) vs. 70.0% (placebo) (p=0.037).
LDN was not effective for inducing remission in CD and showed no benefit over placebo.
That’s an understatement. If clinical remission occurred in far more placebo patients, LDN was harmful.
 
While LDN appears to be effective in other chronic inflammatory diseases such as multiple sclerosis and fibromyalgia [10, 11, 13, 26], its utility in IBD remains unproven. Our results indicate that LDN did not induce remission in patients with CD at least the studied dose and treatment period.
10 is a protocol for an FM study.

11 is a retrospective chart review for MS.

13 is a paper by Jarred Younger on LDN for pain that says that it was highly experimental at the time in 2014z

26 is a Norwegian register review as a quasi-experimental study that showed no signs of benefit for MS.

They have also missed the recent null results in FM and LC.

Overall, I think the authors are pretty biased in their reporting of the literature, and have made some inexcusable errors in how other results are described.
 
What does the p-value have to do with it?
It gives a sense of the probability that the difference in observations was due to random chance.

I was just imagining the inverse case, where a trial comes back positive with a p-value just under 0.05 — we wouldn't be eager to accept this, especially if many comparisons where made.

On the other hand, if it came back positive, it would definitely be interpreted positively, so maybe if the difference points the other direction, it would only be consistent to interpret it as harmful.
 
It gives a sense of the probability that the difference in observations was due to random chance.

I was just imagining the inverse case, where a trial comes back positive with a p-value just under 0.05 — we wouldn't be eager to accept this, especially if many comparisons where made.

On the other hand, if it came back positive, it would definitely be interpreted positively, so maybe if the difference points the other direction, it would only be consistent to interpret it as harmful.
I think the p-value is the likelihood of getting some data given the hypothesis you’re testing. What we would want to know is the likelihood of the hypothesis being true given the data we have.

So the p-value isn’t of much use. There’s an excellent writeup of some of the problems here (look under «problem 3»).

But I agree that we need equality. If they would have acknowledged a positive result with the same p-value, they should have done the same for a negative result.
 
I think the p-value is the likelihood of getting some data given the hypothesis you’re testing. What we would want to know is the likelihood of the hypothesis being true given the data we have.
Yes you are right and I was being a bit sloppy in my interpretation and trying to weasel out using "a sense of".

I assume they did double-sided testing with the null hypothesis being equal means. So the problem is symmetrical.

So we should also apply the same standards as for positive results, that is, that it's probably just null because there's no correction for multiple comparison (I just had a quick look, maybe I missed it).
 
Yes you are right and I was being a bit sloppy in my interpretation and trying to weasel out using "a sense of".

I assume they did double-sided testing with the null hypothesis being equal means. So the problem is symmetrical.

So we should also apply the same standards as for positive results, that is, that it's probably just null because there's no correction for multiple comparison (I just had a quick look, maybe I missed it).
I think we all do that from time to time! It’s usually not an issue with clearly positive results because differences between the groups are so large.

You’re right that it was two-sided, and I can’t find anything about corrections for multiple comparisons in the statistics section either.
Categorical outcomes (remission/response rates) were compared between groups using χ2 test. Continuous outcomes (e.g., change in SES-CD, HBI, biomarkers) were compared using Mann–Whitney U test if non-parametric. A two-sided P < 0.05 was considered statistically significant
The discussion only mentions the positive difference for fatigue and does not mention the negative difference for clinical remission. That’s another clear sign of bias from the authors.
 
Agree with the above comments. You probably know this, but for anyone else, being really pedantic, I'd combine fst's description and the linked blog into "the p-value is the probability of obtaining the outcome you got (or an outcome 'more extreme') if reality matched the null hypothesis".

The fact that it's only a statement about the likelihood of the measured outcome in the null hypothesis world (rather than being about the hypothesis you actually want to prove) is a big tripping block that causes misuse. People will rule out the null hypothesis by reporting a small p-value and then interpret that as if it means their preferred hypothesis is correct, when in fact there may be a dozen other perfectly reasonable hypothesis that they would need to rule out before they can say that. (E.g. you could imagine an experiment showing that horse that does math gets the right answer more often than expected due to chance alone. The experiment may report they had a p-value of 0.001 or whatever, but that's the probability of the horse getting that number of math problems right if it were guessing randomly, and the hypothesis you actually probably want to rule out is that the horse isn't getting subtle cues from the horse trainer.) Wikipedia covers this pretty well now.

On the other hand, p-values are useful if used properly in the situations they were meant for, as in gwas like DecodeME. So using p-values properly comes down to designing experiment where that the null hypothesis is the main alternative possibility that needs to be ruled out. DecodeME tried very hard to get an unbiased sample. And in terms of p-values, you could say that they did this so that when a location in the genome turned out to have a small p-value we can be pretty hopeful that's actually due to our preferred hypothesis (that that location is associated with ME/CFS) and not just due to some confounder (which would be one of those hypothesis that's neither the null hypothesis nor our preferred hypothesis).
 
Back
Top Bottom