Protocol for studying ME/CFS and Long Covid

bunkrocket

New Member
Hello,

First time posting. I wasn't sure where to put this thread.

Browsing this forum has really shown me how flawed so many of the study designs have been. I was wondering if it'd be helpful to outline what a good clinical protocol should actually look like for studying ME/CFS and Long covid?

Some examples : what does a proper control sample look like?
What defitions and criteria should be used (Canadian Consensus Criteria, DSQ-PEM, FUNCAP)

If this has already been discussed and I've missed it then please ignore!
 
I was wondering if it'd be helpful to outline what a good clinical protocol should actually look like for studying ME/CFS and Long covid?

I think the answer is that it will depend entirely on what question one is trying to answer. Appropriate control groups and inclusion criteria will depend on what you want to explain or treat. Looking at wider or narrower groups is fine if there are good reasons to do so.

Studies tend not to stand or fall on whether or not they conform to a norm. They depend on getting the design detail right for the task. PACE was not useless because it used Oxford criteria. It was useless because of the open design with subjective outcomes.
 
Welcome to the forum, @bunkrocket.

I agree that a list of criteria to satisfy may not be that practical. The other direction might be more useful: a list of what not to do, based on common problems in ME/CFS literature. Maybe someone else here knows of a good overview for that. The only work I'm aware of that comes close is this excellent satirical piece: A Field Guide to Conducting Biopsychosocial Research in CFS/ME. This is specifically about BPS, but methodological issues undermining conclusions is not limited to researchers holding these views. Some biomedical publications are just as guilty. More than once, this forum interprets data from biomedical papers as evidence for the opposite of what the authors write in their conclusions.

I originally continued my post with the following "I would expect that researchers are aware of all of this methodology stuff from uni courses. Why problems persist, I don't know. I'm not sure if any list of do's and don'ts would influence those that commit these errors", but on second thought, I think it could be valuable for younger researchers or those not too familiar with ME/CFS to be aware of common methodological issues in the field. It could be a valuable resource for patient advocates and newcomers interested in the flaws of the research that is claimed to support biopsychosocial ideas.

Again, if anyone does have a reference to such a list, I'd like to read it!

I also want to mention an example of ME/CFS trial methodology that is positively received here: Fluge and Mella's work.
 
PACE was not useless because it used Oxford criteria. It was useless because of the open design with subjective outcomes.
That's an interesting statement. How would you quantify the impact of using the Oxford criteria on the relevance of results? Would it not by default make the evidence of low quality through population indirectness?
 
The Oxford criteria cover most people with MECFS so any valid results from PACE would apply. If PACE showed evidence for a need to subset that might alter application but we dont have that. I personally think the indirectness argument is a distraction. All we need to know is that PACE did not show meaningful evidence of benefit across the cohort.
 
That's an interesting statement. How would you quantify the impact of using the Oxford criteria on the relevance of results? Would it not by default make the evidence of low quality through population indirectness?
The Oxford criteria cover most people with MECFS so any valid results from PACE would apply. If PACE showed evidence for a need to subset that might alter application but we dont have that. I personally think the indirectness argument is a distraction. All we need to know is that PACE did not show meaningful evidence of benefit across the cohort.
If the intervention is only effective for one subgroup, doing a study with a population that also includes other subgroups might obscure the real effect on the first subgroup. An exaggerated example is that a trial of radiation therapy for a group of patients that are coughing would get a negative result even if the therapy is effective for the subgroup that has lung cancer.

Meaning that a null result in isolation might not mean that it doesn’t work for some. Although you can’t claim that it works for some unless you get a positive result in a subsequent trial that only includes the subgroup of interest.

At the same time, if the intervention actually doesn’t work for anyone in the larger group, by definition it won’t work for anyone in the subgroups.

So when you have a null result in a broad population, the question becomes if it’s because the intervention doesn’t work for any, or if it’s because it only works for too few to show up in the statistical analyses.
 
If the intervention is only effective for one subgroup, doing a study with a population that also includes other subgroups might obscure the real effect on the first subgroup.

Yes, but this will always apply, ultimately down to the individual, on which you cannot get statistical data. A wider set of criteria is less of an issue than a too narrow one. For rheumatoid it turns out that you get no response to rituximab if there is no sign of abnormal Ig production (as you might expect). But you start with the wider RA criteria and then do some subgroup work once you have a basic result.
 
So when you have a null result in a broad population, the question becomes if it’s because the intervention doesn’t work for any, or if it’s because it only works for too few to show up in the statistical analyses.

In this situation you have to rethink your rationale because you are going to be chasing a small minority and your original grouping was clearly not right for the context. Basicallly, you start again.

The problem with invoking indirectness is that it provides a tacit recognition that CBT and GET might be useful for some other people. We have no reason to think that. It also makes people focus on PEM as a discriminator and I don't personally see that as robust. You end up going off on tangents.
 
Yes, but this will always apply, ultimately down to the individual, on which you cannot get statistical data.
Sure, I agree with that.
A wider set of criteria is less of an issue than a too narrow one.
Why is that? (I might have answered below)
For rheumatoid it turns out that you get no response to rituximab if there is no sign of abnormal Ig production (as you might expect). But you start with the wider RA criteria and then do some subgroup work once you have a basic result.
Don’t you run the risk of overlooking the rarer cases then?

I can understand if this is the most practical way to do it, and you can also argue that your first priority is to find treatments that will help the majority (if that exists), so starting broad is the best approach. So I’m just trying to understand the drawbacks.
In this situation you have to rethink your rationale because you are going to be chasing a small minority and your original grouping was clearly not right for the context. Basicallly, you start again.
I agree. It would for instance indicate that chronic fatigue and ME/CFS is not part of the same overarching category.
The problem with invoking indirectness is that it provides a tacit recognition that CBT and GET might be useful for some other people. We have no reason to think that.
I agree, and I’m frequently frustrated by advocates and professionals that keep saying that it might be useful for people with just fatigue and so on, just not for ME/CFS.
It also makes people focus on PEM as a discriminator and I don't personally see that as robust. You end up going off on tangents.
As in PEM will always mean ME/CFS, so nobody else experiences PEM?

I agree that that argument gets us into some weird situations, and it’s not made better by the false positives of DSQ and in general PEM being classified as any post-exertional symptoms by some. You’ll have a very hard time drawing a clear line anywhere.
 
Why is that?

A wider set of criteria will at least include all the people you later decide you are interested in (maybe CCC ME/CFS). Too narrow a set of criteria will mean you do not even have a chance of studying a proportion of your favoured group. If the treatment has adverse effects on those you will not pick it up in the narrow study.

The basic rule is that a statistical conclusion will apply to subsets of a set studied unless there are reasons to think otherwise. The reverse is not true. A statistical conclusion does not apply to a superset. Take a set of people under five feet high who can walk through a low door. That applies to a subset only four feet high but not to a superset that includes people up to six feet high.

There are arguments for choosing narrow criteria, and hence homogeneity, in testing initially for an effect if you thik it it may be hard to get statistical significance from a more heterogeneous group. But that is a different issue and applying the result to a wider group later (as you may do) is based on inference rather than evidence.
 
The problem with invoking indirectness is that it provides a tacit recognition that CBT and GET might be useful for some other people. We have no reason to think that. It also makes people focus on PEM as a discriminator and I don't personally see that as robust. You end up going off on tangents.
I agree, and I’m frequently frustrated by advocates and professionals that keep saying that it might be useful for people with just fatigue and so on, just not for ME/CFS.
These arguments are new to me and conflict with my current understanding. @Jonathan Edwards, is your opinion that CBT does not work for depression? A lot of stuff can be interpreted as "chronic fatigue". I would have a hard time arguing that CBT or GET are ineffective for the majority of them. Could you expand on your frustration, @Utsikt?

A wider set of criteria will at least include all the people you later decide you are interested in (maybe CCC ME/CFS).
It may also be so wide that a sample may not include the subgroup you're actually interested in. This study reports Oxford criteria includes something like 15 times more people than CCC.
 
These arguments are new to me and conflict with my current understanding. @Jonathan Edwards, is your opinion that CBT does not work for depression? A lot of stuff can be interpreted as "chronic fatigue". I would have a hard time arguing that CBT or GET are ineffective for the majority of them.

OK, but I have been repeating these arguments for 5-7 years now and made the central point in my oral submission to UK NICE in 2019/20. I am aware that there has been a wide assumption that the problem with PACE lies with inclusion criteria and that this affected the NICE decision. As I have said, there is technically no reason why the PACE team should not have chosen Oxford and if they had valid results those would apply unless there was other evidence. The NICE decision was initially influenced by 'indirectness'. However, at the Round Table meeting with objectors Peter Barry gave a re-analysis in which the decision that CBT and GET were not cost effective was not altered by removing the indirectness calculation.

I am not sure what you mean by 'a lot of stuff' being interpreted as chronic fatigue. Ill people all deserve listening to and caring for. ME/CFS dooes seem to be a useful subset of those whose symptoms include 'fatigue' as understood (poorly) in the medical world. If PACE showed that CBT and GET helped people with fatigue generally and found no evidence for people with CCC criteria being different in the regard (they say they did not find a difference) then the therapies would be recommendable.

Why would you find it hard arguing that CBT and GET are ineffective for fatigue? I cannot see any reason why they should be nd we do not have any decent trials to show they are. PACE does not show they are.

When I pointed out the serious flaws in PACE to the NICE committee someone asked me to confirm that I was criticising methodology for CBT evidence specifically for ME/CFS rather than for CBT in general. I said that my analysis had been restricted to ME/CFS but that having seen the way psychiatrists defended PACE I thought it likely that the evidence forCBT was probably just as bad in other areas. I have no idea whether CBT helps depression and I doubt anybody else does. I would bet the trials are as bad as PACE.

I am very happy to believe that talking to depressed people helps them cope but I am pretty sceptical that the methodology known as 'CBT' has anything specific to offer.

And depression is not fatigue. So all in all I see no good argument for suggesting that CBT might be useful for 'non-ME/CFS fatigue', whatever that covers.
 
It may also be so wide that a sample may not include the subgroup you're actually interested in. This study reports Oxford criteria includes something like 15 times more people than CCC.

That is a theoretical possibility but my memory is that a substantial proportion of PACE subjects satisfied ME/CFS criteria. Moreover, even though the PACE authors did not call it ME/CFS their rationale was aimed at changing cognitions of people who believed that exercise made them worse - and those are surely likely to be those with ME/CFS. They ended up showing that despite maybe shifting what subjects thought a bit CBT and GET did not actually produce a credible reduction in disability, disproving their theory.
 
The basic rule is that a statistical conclusion will apply to subsets of a set studied unless there are reasons to think otherwise.
As in a rule of thumb?

My coughing and lung cancer example shows that if the subgroup is small enough, you won’t see it in the group level analyses.

I’m willing to accept that in general that kind of situation is very unlikely to happen, at least not in serious medical research. But we’ve seen plenty of bad biomed research. Although you could argue that the chance of any of them finding an effective treatment for anything is so low that we can ignore it.
I would have a hard time arguing that CBT or GET are ineffective for the majority of them. Could you expand on your frustration, @Utsikt?
My frustration lies with the terrible methodology and equally terrible assessments of the methodology in meta analyses.

If you can show me any large analyses on depression of chronic fatigue that demonstrates a positive result in the absence of significant bias, I would be at least partially wrong.
 
As in a rule of thumb?

No, as in Bayes's theorem. Without prior statistical weighing expectation for the subset is that for the set.
My coughing and lung cancer example shows that if the subgroup is small enough, you won’t see it in the group level analyses.

But we aren't erxp-ecting the subset to be so small, as indicated above.

I’m willing to accept that in general that kind of situation is very unlikely to happen, at least not in serious medical research. But we’ve seen plenty of bad biomed research.

All we can work with is probabilities. By Bayes's theorem you have to be allowed to draw conclusions from wider sets to smaller. You aways are with individuals. But, like all probabilities in Bayes's theorem, they can change with new information.
 
Thank you for your responses.
I am very happy to believe that talking to depressed people helps them cope but I am pretty sceptical that the methodology known as 'CBT' has anything specific to offer.
That is my view as well. That is why I disagreed with your earlier statement of:
The problem with invoking indirectness is that it provides a tacit recognition that CBT and GET might be useful for some other people.

I do see that the "by default" in this statement was too bold:
How would you quantify the impact of using the Oxford criteria on the relevance of results? Would it not by default make the evidence of low quality through population indirectness?

I would still say "using the Oxford criteria means there is a potential risk of population indirectness" (not talking about any specific study here). And if we agree that talking to a depressed patient who somehow gets the label "chronic fatigue" (or exercising a deconditioned patient who gets that label) can improve their condition, there is a potential for population indirectness to lead to spurious results.

I agree, and I’m frequently frustrated by advocates and professionals that keep saying that it might be useful for people with just fatigue and so on, just not for ME/CFS.
Would you agree if instead of "people with just fatigue" is changed to "people labeled with chronic fatigue but don't have PEM"?

That is a theoretical possibility but my memory is that a substantial proportion of PACE subjects satisfied ME/CFS criteria.
If that 15x stat is correct, very skewed samples are not just a theoretical possibility but very likely. But I think I didn't communicate well in my last message I was talking about the indirectness argument in general, not applied to the PACE trial.
 
No, as in Bayes's theorem. Without prior statistical weighing expectation for the subset is that for the set.
So you’re saying that the probability that the subset has the same properties as the whole set with regards to the intervention that was tested, is by default greater than the probability of the opposite being true?

I don’t have the bandwidth to write out the equation now, and I’m not even sure which numbers you’d use for the different parts.
But we aren't erxp-ecting the subset to be so small, as indicated above.
I was talking in general terms, not about PACE specifically.
Would you agree if instead of "people with just fatigue" is changed to "people labeled with chronic fatigue but don't have PEM"?
You can use whatever group you want, I don’t think there is any group where CBT has been proven to have a positive effect.
 
Back
Top Bottom