Artificial intelligence in medicine and science

"I Was an Oncologist. When I Got Sick, I Did What Doctors Warn Patients Never to do"

"I have hypermobile Ehlers-Danlos syndrome, a connective tissue disorder that affects nearly every part of my body. My joints dislocate with alarming ease. My gastrointestinal tract operates by its own rules. My autonomic nervous system misfires constantly. I have chronic pain, episodes of exhaustion, medication sensitivities and symptoms that refuse to stay inside one specialty. Managing a disease like mine requires a cardiologist, gastroenterologist, neurologist, rheumatologist, pain specialist, rehabilitation physician and a primary doctor all communicating with one another in ways modern medicine rarely allows."

"Eventually, I did what patients are warned not to do. But I was also a physician, trained to stay with difficult medical problems until they made sense. I immersed myself in everything I could learn about each condition and how it affected every part of my body.

Then I wanted another mind on the case, the way physicians bring cases to a tumor board.

I turned to artificial intelligence.

Even now, writing that feels like a professional betrayal. As a physician, I understand exactly how dangerous AI can be in medicine. It can hallucinate, provide dangerously incorrect information and miss diagnoses. It lacks the accountability clinical care demands. AI should never replace physicians."

I began hearing similar stories from physicians treating other poorly understood chronic illnesses. A leading mast cell activation syndrome expert told me about a woman whose neuropsychiatric symptoms went undiagnosed for six years. Out of curiosity, he uploaded years of hospital and clinical records into ChatGPT. ChatGPT identified MCAS almost immediately and laid out a rationale that aligned with the diagnosis the specialist later reached after three weeks of evaluation.

There is something else AI offers that medicine often does not".
They could have got the same answer on social media where hEDS and MCAS are commonly linked.
 
I’ve seen some really interesting work with adversarial iteration until agreement is reached. Sounds like what you’re doing?

No, I wouldn't necessarily make sure that there is an agreement. Basically I start with -for example- several findings from ME/CFS studies (e.g. GWAS, replicated findings etc) and then I ask a number of LLMs (e.g. 3) to suggest a causal hypothesis.

2) Next step is to present to each LLM the hypotheses of the other two and rank it.

3) A PageRank algorithm (this is an example) gives a ranking of the hypotheses and how much each LLM agrees that another hypothesis is better.

4) An "orchestrator" LLM looks at all hypotheses, identifies where all LLMs agree, where they disagree, whether any factual discrepancies were found and whether novel information deserving further investigation has been identified.

It works surprisingly well in my opinion.
 
Thanks for explaining @mariovitali an interesting approach! If you’ve written more on this anywhere please do share any links.

On the topic of this thread IMHO the solution in the healthcare industry is for those in it to provide better support. Then fewer people will turn to unsuitable alternatives.

Meanwhile AI is achieving real results in STEM, here’s yet another

They’re not magic answer machines and selling or using them as such is dangerous and likely starting to hit walls in terms of financial viability https://arstechnica.com/ai/2026/06/...sers-react-to-new-usage-based-pricing-system/

But these technologies do and will continue to have uses and achieve breakthroughs in the sciences including medicine. To me a lot of discourse seems to end up in unhelpful battles over things being entirely good or bad or focus on the morals of the technology rather than the people, business practices and regulation. We need better education and better informed decision making on how we want to use the technology as it evolves.
 
So I just learned about Lean, which is a programming language that essentially allows you to check if a math solution is logically valid or not.

Which means that for all maths problems AI tries to solve, they will be able to check if the solution they are working on actually is mathematically sound while they are working on the problem.

This makes it far easier to solve math problems than problems where you don’t have a way to check if your solution actually makes sense. The new math proofs are exciting, but still a substantially different use case compared to medicine.
 
Yes exactly, AI seems promising in domains where we can automatically check correctness of the final answer. In math, my understanding is it's operating as an idea generating engine, and there's a second piece of the system checking the ideas the model comes up with and stitching them together (presumably using Lean as you suggest). We should be cautious about extrapolating AI being good at advanced math to anything else.

Ironically, I am coming to think this property (which is basically falsifiability I guess) is also why humans have been able to solve such hard problems in math compared to certain other fields, where we just seem to go in circles without ever learning from our mistakes...
 
Nature: AI agents are checking the scientific literature — and spotting decades-old errors

Many researchers already dedicate their time to spotting errors in papers and use tools to check certain facets of papers. But one advantage of using AI is the speed at which it can scan scientific databases and literature compared to humans, says James Zou, a computer scientist at Stanford University, California. “The biggest difference is to be able to do this at a scale that was not possible before.”

In a study posted on the preprint server arXiv1, Zou and his colleagues used an ‘AI checker’ to scan papers published at NeurIPS — a prestigious annual AI research conference — for errors. Their tool found that errors in papers rose from 3.8 in 2021 to 5.9 in 2025 — an increase of 55%.

“These are papers that have been published, so they’re sort of taken as the foundational knowledge for the next generation of research,” says Zou. “If there are mistakes in these foundations, this can propagate and make the follow-on research shakier,” he adds.

The analysis focused on ‘objective’ errors, such as those in formulae, calculations and figures, and excluded subjective mistakes about data interpretation and novelty.

@rvallee
 
Another good example of a positive use case, making gene editing safer.

A couple of decades after the discovery of systems that could selectively target DNA, we’re starting to see the first therapies based on gene editing. One challenge these developments have faced is safety. While we can make them pretty specific to the gene we want edited, the human genome is very large, and even rare DNA sequences can appear a couple of times by chance.

As a result, all the original gene-editing systems had known rates of what are called off-target effects, in which they simply edit the wrong sequence. This may be a low-probability event, but edit enough cells—and therapies generally have to edit many—and errors become inevitable.

A lot of effort has gone into finding ways to minimize or eliminate off-target edits. In a recent issue of Nature, researchers described modifying the AI protein-folding software AlphaFold to help identify key areas of gene-editing proteins responsible for off-target effects. Those areas were then modified to reduce the problems.
the researchers also suggest that the approach would be useful more generally for fine-tuning protein-DNA interactions, which could have applications far beyond gene editing.

Precise DNA base editing using AlphaFold3-based contact modelling, 2026, Meng et al

Meng, Haowei; Lei, Zhixin; Yan, Yongchang; Wang, Liren; Zhang, Sihan; Rao, Xichen; Shao, Chuyun; Zhang, Xiaoting; Chen, Ke; Yang, Lei; Liu, Rongrong; Yang, Gaohui; Shen, Ruoyu; Gu, Ruichu; Wang, Xinyan; Wang, Yiya; Lu, Suiru; Lv, Zhicong; He, Bo; Wen, Han; Li, Dali; Yi, Chengqi

Abstract
Achieving high specificity in biochemical transformations is crucial for research and therapeutics. This is particularly important for genome editing, where enhancing tool specificity ensures effective and precise editing outcomes. Current strategies are constrained by activity-specificity trade-offs, high labour intensity and low success rates. Here we present ContactSeek, an artificial-intelligence-driven framework that uses AlphaFold3 (AF3)-predicted contact probability5 to improve the specificity of genome editors. Using Cas9–TadA adenine base editors as a demonstration, we mapped their genome-wide off-targets and fed the off-target DNA sequences to AF3. Among AF3 outputs, we found that contact probability was more sensitive than predicted three-dimensional structures for detecting differential interactions between on- and off-target complexes. Correlating contact probability with sequencing-based off-target signals, ContactSeek identified and ranked consensus contact regions, which are neighbouring Cas residues with consistent contact changes to DNA/guide RNA, and pinpointed specificity-determining residues within them. ContactSeek can also be applied modularly and identified key residues in the TadA8e deaminase. Targeted amplicon sequencing, genome-wide profiling, R-loop assay and RNA-sequencing together confirmed the greatly enhanced specificity; our best variant, combining two mutations of Cas9 and TadA8e, outperformed several known high-fidelity adenine base editors. ContactSeek is also generalized to Cas12a-based cytosine base editors. Collectively, our framework represents an AF3-driven model tailored for specificity improvement, establishing a paradigm for improving the precision of genome editing tools through the integration of structural and functional dimensions.

Web | DOI | PDF | Nature
 
Anthropic (Claude company) and other AI companies have started implementing hidden watermarks in their AI-generated text, to comply with EU regulations. It adds a slight bias in which words are picked, allowing someone to check if the text was likely written by one of these company's AIs.

Anthropic: "How Claude’s text watermark works"
  • We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;
  • The difference between watermarked and un-watermarked text will not be distinguishable to readers;
  • Nothing is added to the text and there are no hidden characters;
  • Watermarking doesn’t require extra tokens, and will not be more expensive;
  • Watermarking carries no identifying information and can’t be traced to a specific person, organization, or chat;
  • Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.
How do I check if a piece of text was written by Claude?
We will soon be offering a watermark detection API. We’re in the process of working out the details of its implementation.
 
Back
Top Bottom