Signals Inbox·August 30, 2026·MedTech

What are the latest AI medical breakthroughs?

The latest AI medical breakthroughs are no longer benchmark wins: AI is now catching cancers missed in routine care, improving specialist diagnoses, removing large parts of screening workloads, guiding surgeons live and pushing an AI-originated drug into Phase 3.

We track MedTech daily. Want the market signals in your inbox?

Send me the signals
Summary

The biggest AI medical breakthroughs right now are AI-assisted specialist diagnosis, late-stage AI drug development, partially autonomous medical imaging, AI systems catching findings missed in routine care and the first live uses of surgical AI.

The important shift is the evidence. Medical AI is finally being tested prospectively and in randomized trials involving tens of thousands of real patients, which makes it much easier to separate genuinely useful systems from impressive demonstrations.

Some of the strongest results come from narrow systems rather than general AI doctors. Retina4IRD improved specialist diagnostic accuracy by more than 20 percentage points, mammography AI cut radiologist workload by 63.6%, and LiON found malignant lesions after the normal radiology process had missed them.

AI drug discovery may ultimately have the largest impact. Rentosertib reaching Phase 3 means an AI-originated molecule has survived several years of biological, safety and human-efficacy tests rather than simply looking promising on a computer.

The less comfortable finding is that good AI does not automatically improve care. Recent trials show that hospital bottlenecks, clinician adoption and downstream capacity can erase much of the benefit of a technically strong system.

100+ new signals every week · 50+ markets · updated daily

Need to know what's hot?We can send you all the signals

Send me the signals Delivered straight to your inbox

Q1What actually deserves to be called an AI medical breakthrough today?

Today, the strongest AI medical breakthroughs are the ones that have crossed into real clinical testing, changed what doctors can do, or pushed an AI-created treatment much further through human trials.

That is a stricter bar than the one often used in AI news. A model beating another model on a medical benchmark can be interesting research, but patients never interact with benchmarks. We care much more when AI is tested prospectively in hospitals, compared with doctors in a randomized trial, used during an actual operation, or attached to a drug that survives human testing.

That stricter definition changes the list quite a lot. Some of the most important advances currently come from fairly narrow systems. LiON reads liver CT scans as an extra pair of eyes. Retina4IRD helps retinal specialists narrow down the genetic cause of rare eye diseases. A partially autonomous mammography system can remove many low-risk exams from the radiologist's queue. Rentosertib is an AI-originated drug that has reached Phase 3 development. UCLH has now used AI on a live surgical video feed during a human brain operation.

These are more important than most "AI doctor" demonstrations because they have crossed harder boundaries between technical performance and actual medicine.

AI medical breakthroughs with the strongest current evidence

AI medical advance What changed Evidence level How big is it?
AI liver-cancer reading Cancers missed on the first read were found in routine care Prospective clinical study Strong
AI-assisted retinal diagnosis Specialists became much more accurate Randomized clinical trial Strong
Partially autonomous mammography AI removed much of the reading workload while detecting more cancers Prospective clinical trial Strong
AI-designed drug development An AI-originated drug reached Phase 3 Human Phase 2 completed Potentially huge
Real-time surgical AI AI analysed live video during human brain surgery First patient in clinical trial Very early but genuinely new
Medical LLM copilots AI can support diagnosis and documentation Prospective and randomized studies Useful, but still mixed

Q2Is medical AI finally being tested on real patients?

Yes, medical AI is finally getting the kind of prospective and randomized testing that was missing for years, although most approved systems still have much weaker clinical evidence.

The gap is easy to see in the FDA record. A recent analysis counted 1,430 AI or machine-learning medical-device authorizations through the end of 2025, up from an average of fewer than two authorizations per year during 1995-2014 to 264 per year during 2023-2025. More than three-quarters were radiology devices.

The evidence grew much more slowly. A JAMA Network Open review examined 717 FDA-authorized radiology AI devices for which submission documents were available. Only 33, around 5%, had undergone prospective testing. Just 15 combined prospective and clinical testing, and six included prospective testing, clinical testing and a human operator.

That makes the newest clinical studies unusually valuable. We now have AI trials involving 31,301 women receiving mammograms, 9,691 primary-care patients treated with or without an LLM copilot, 93,326 chest X-rays in a randomized lung-cancer pathway study, and prospective AI liver analysis across 10,333 patients. Those sample sizes put some medical AI questions on much firmer ground than they were even recently.

Regulators are reacting to the same shift. The FDA has just opened a dedicated discussion on generative-AI medical devices covering premarket testing, risk, postmarket monitoring, foundation models and agentic systems. That would have been a fairly theoretical exercise a few years ago. These days, regulators are preparing for products that can generate and reason rather than simply classify an image.

The evidence problem certainly has not disappeared, but the best medical AI research now looks much more like clinical medicine and much less like a computer-science benchmark contest.

Q3Is AI now catching cancers that radiologists miss?

Yes, AI is currently finding some cancers that radiologists initially miss, and the latest large prospective liver study gives us unusually concrete evidence.

The LiON system analyses contrast-enhanced CT scans for liver malignancies. Researchers first trained it on 6,443 patients and retrospectively tested it across another 22,251. The more interesting test came when LiON was added to the normal hospital workflow as an extra reader for 10,333 patients.

According to the recent Nature Medicine study, AI-human review found 51 lesions that had been overlooked in the original interpretation. Fifteen turned out to be malignant. Those findings led radiologists to amend 37 reports and sent 22 cases for multidisciplinary review.

Put that into a more intuitive scale: across the prospective cohort, the system helped surface roughly one previously overlooked malignancy for every 689 patients processed. That number sounds small until we remember what is being counted: cancers that had already passed through the normal reading process.

LiON's prospective AUC was 0.952, close to its 0.975 result in the much larger retrospective validation. Performance also held up reasonably well in patients with cirrhosis and fatty liver, two groups where liver imaging can be harder to interpret.

LiON has not proved better cancer outcomes. The prospective study had no randomized control group, and it did not measure survival. But it does give us strong evidence that AI can work as a diagnostic safety net inside a real radiology department and retrieve malignant findings that humans have already overlooked.

We track MedTech daily. Want the market signals in your inbox?

Send me the signals

Q4Can AI safely take over part of breast cancer screening?

Yes, AI can already take over a large part of routine mammography reading, but the trade-off is more women being called back for additional checks.

A prospective study of 31,301 women compared conventional double reading with a partially autonomous AI strategy. Low-risk mammograms identified by the AI were treated as normal without the usual human reading process, while higher-risk cases continued through radiologist review with AI support.

Radiologist workload fell by 63.6%.

That is a huge operational change. In practical terms, the AI strategy removed almost two-thirds of the normal reading work while cancer detection increased from 6.3 to 7.3 cancers per 1,000 screenings. The researchers measured a 15.2% higher cancer-detection rate.

The cost was a 14.8% increase in recalls, meaning more women were asked to return for additional assessment. The strategy therefore missed the study's predefined non-inferiority target for recall rates.

The result is more interesting than simply saying AI reads mammograms well. Health systems are beginning to test whether AI can remove humans from low-risk cases and concentrate radiologists on the harder ones. That changes staffing and screening capacity, especially in places where radiologist shortages limit how many women can be screened.

The recall problem is still meaningful. If autonomous screening saves thousands of radiologist reads while generating too many unnecessary investigations, the economics and patient experience change with it. The workload advantage is already very large. The remaining argument is how much extra recall health systems are willing to accept for that gain.

Q5Can AI actually make specialist doctors much better at diagnosis?

Yes, AI can already make experienced specialists substantially more accurate on some difficult diagnoses, and inherited retinal disease gives us one of the clearest randomized examples.

Retina4IRD helps doctors predict the genetic cause of inherited retinal diseases from retinal photographs, optical coherence tomography scans and clinical information. The system was developed using genetically confirmed cases from 1,843 patients across China, South Korea and Poland.

Researchers then ran a randomized clinical trial involving 300 people with suspected inherited retinal disease. Of the 295 patients included in the final analysis, specialists using Retina4IRD put the correct genetic category within their top five possibilities 88.5% of the time. Specialists working without the AI reached 67.3%.

That 21.2-percentage-point gap is large. Roughly speaking, adding the AI produced one additional correct top-five genetic hypothesis for every five patients assessed.

The advantage also appeared when doctors had fewer guesses. Top-one accuracy rose from 22.4% without AI to 37.8% with it, while top-four accuracy went from 53.1% to 81.8%. The researchers also found better downstream management decisions in the AI-assisted group.

Experienced doctors still had plenty of room to improve. Rare retinal diseases can involve many overlapping visual patterns and genes, so even specialists face a long tail of combinations they do not see often.

This is probably one of the most credible ways medical AI will spread in the near term: a specialist remains responsible for the diagnosis, while the machine gives that specialist access to a much broader memory of unusual disease patterns.

Q6Can AI diagnose rare diseases better than experienced doctors?

On difficult rare-disease cases, DeepRare can currently outperform experienced physicians, although we still do not know whether it can shorten the real diagnostic journey for patients.

Rare diseases are almost designed to expose the limits of human memory. More than 7,000 rare diseases are recognized, and many doctors may never encounter a specific one during their careers. Patients commonly spend years moving between specialists before receiving a diagnosis.

DeepRare approaches that problem with a multi-agent system connected to more than 40 medical tools and knowledge sources. The researchers evaluated it across nine datasets covering 2,919 diseases, 14 specialties and patients from Asia, North America and Europe.

On phenotype-based tasks, DeepRare achieved average top-one recall of 57.18%, beating the next-best method by 23.79 percentage points. In another test involving 168 cases with genetic information, it reached 69.1% compared with 55.9% for Exomiser, a widely used disease-prioritization system.

The human comparison is even more striking. On a hospital case set, DeepRare placed the correct diagnosis first more often than five experienced physicians and also performed better when allowed several candidate diagnoses.

"Better" still needs context. Those doctors were solving prepared cases rather than taking responsibility for patients from the first consultation through testing and treatment. Real diagnostic medicine involves deciding which information to collect in the first place, noticing when a patient explains something poorly, ordering tests and revising hypotheses as new evidence appears.

Still, rare-disease diagnosis looks unusually well suited to AI. The system can search an enormous body of obscure medical knowledge within seconds, which attacks one of the actual bottlenecks in rare disease rather than merely automating paperwork.

100+ new signals every week · 50+ markets · updated daily

Need to know what's hot?We can send you all the signals

Send me the signals Delivered straight to your inbox
Market Signals

Q7Are AI medical copilots actually improving patient care today?

Medical AI copilots are useful today, but the best clinical studies still do not show a reliable improvement in patient outcomes.

That is an important correction to the enormous gains seen in medical LLM benchmarks. Google's multimodal AMIE, for example, can conduct simulated consultations using patient conversation, skin photographs, ECGs and clinical documents. In a randomized evaluation of 105 simulated telehealth cases, specialist reviewers rated AMIE similar to or better than primary-care physicians across most measured dimensions.

Real clinics are proving tougher.

A randomized primary-care trial across 16 facilities in Kenya involved 9,691 patients and 103 clinical officers. Clinicians using an LLM assistant produced better documentation and more complete plans, but 14-day treatment failure was 2.2% with AI and 2.0% without it. The difference was not statistically significant.

A very recent Nature Medicine study gives us another reality check. SHAKED, an emergency-department decision-support system built on several LLMs, was prospectively evaluated across 1,138 patients. Expert reviewers considered 99 of 100 sampled AI outputs clinically appropriate, and researchers detected no AI-related adverse events.

Doctors gradually stopped using it anyway.

SHAKED usage fell from 68% to 30% over four weeks. Emergency-department stays remained 4.9 hours in both groups, and the roughly nine-minute reduction in consultation-cycle time was not statistically significant.

That is more revealing than another physician-exam victory. Modern medical LLMs can already produce useful clinical reasoning. The harder challenge now is getting busy doctors to use that reasoning consistently and then showing that patients actually do better.

Q8Is live AI-assisted surgery already happening?

Yes, live AI-assisted surgery is now happening in humans: UCLH has used AI to analyse a neurosurgeon's video feed during an actual brain-tumour operation.

The case is extremely recent and still only the first patient in a clinical trial, so its scale matters. But technologically, something genuinely new happened.

The patient had an 11-millimetre pituitary tumour close to structures involved in vision. During surgery at the National Hospital for Neurology and Neurosurgery in London, the AI analysed the live endoscopic camera feed and highlighted important anatomy around the base of the brain while the surgeon worked.

The surgeon remained in full control. The AI's job was visual guidance: helping identify structures such as nerves and blood vessels in an area where a mistake of only a few millimetres can cause blindness, stroke or other severe complications.

UCL says the system learned from hundreds of annotated pituitary-surgery videos. Before this clinical trial, the technology had already been evaluated and used as a surgeon-training tool. The jump from recorded video to a live human operation is what makes the latest case interesting.

The tumour was removed and the patient's vision improved, but a single successful operation cannot tell us whether the AI reduces complication rates. The ongoing trial will need to establish that.

We are much closer to AI-enhanced surgeons than autonomous surgeons. In the near term, computers can watch the operative field continuously, recognize structures, track instruments and warn about dangerous areas while the human surgeon handles the operation itself. That is already a meaningful change in what surgical software can do.

Q9Has AI actually created a drug that reached Phase 3?

Yes, AI drug discovery has now produced an AI-originated molecule that reached Phase 3 development, which is far more meaningful than generating promising compounds on a computer.

Rentosertib, developed by Insilico Medicine for idiopathic pulmonary fibrosis, is one of the strongest examples. Insilico used AI in both target discovery and molecule design, identifying TNIK as a target and then developing a small-molecule inhibitor against it.

The drug subsequently made it through preclinical work, Phase 1 testing and a randomized Phase 2a study involving 71 patients.

In that Phase 2a trial, patients receiving the highest tested dose had a mean 98.4 mL increase in forced vital capacity after 12 weeks, while the placebo group declined by 20.3 mL. The raw difference between those groups was 118.7 mL.

The trial was small, so that figure is not settled efficacy. Safety also needs much more testing. What makes rentosertib different from hundreds of "AI discovered a drug" announcements is how far it has travelled through the pipeline.

The company has now started Phase 3 development, with a planned study of roughly 320 participants and a much longer lung-function endpoint.

Drug discovery may ultimately become the largest medical effect of AI because even a modest improvement in the probability of finding successful molecules would compound across thousands of drug programs. Rentosertib currently gives us the clearest test of that thesis.

How far Rentosertib has moved through drug development

Stage What usually kills an AI-drug claim Rentosertib
Target discovery Interesting biology never produces a workable drug Passed
Molecule generation Generated molecule fails preclinical development Passed
Phase 1 Safety or dosing problems stop development Passed
Phase 2 Drug fails to show enough human activity Positive Phase 2a result
Phase 3 Large-scale efficacy remains unproven Now being tested
Approval Drug must still satisfy regulators Not reached

We track MedTech daily. Want the market signals in your inbox?

Send me the signals

Q10Is AI finding new antibiotics humans would probably miss?

Yes, deep-learning systems are now finding unusual antibiotic candidates across chemical spaces far too large for humans to test molecule by molecule.

A recent study targeting drug-resistant Neisseria gonorrhoeae started by physically testing 38,650 small molecules. Researchers used the results to train a graph neural network that could predict which other molecules might stop the bacteria growing.

The trained system then screened around six million compounds virtually.

Only 213 were brought back into the laboratory for physical testing. Eighty-three worked, giving the AI-selected set a 39% hit rate.

The scale of that funnel is worth looking at. The initial experimental screen gave the model tens of thousands of examples; the AI then searched a chemical universe more than 150 times larger and reduced it to a few hundred candidates. More than a third of those candidates showed activity when researchers actually tested them.

Several promising molecules were structurally different from known antibiotics, which is especially valuable in a field where resistance can quickly undermine drugs that attack familiar biological targets. Selected candidates also worked against multidrug-resistant strains and showed activity in laboratory tissue models and mice.

None of this gives doctors a new gonorrhoea drug today. Preclinical antibiotic candidates still face toxicity, dosing and human efficacy testing, and most experimental drugs never reach pharmacies.

The breakthrough here is the search process. AI can narrow millions of possible molecules to a tiny experimental shortlist with a remarkably high hit rate. If that efficiency survives across more diseases, medicinal chemists gain access to parts of chemical space they could never explore manually.

Q11Can AI tell which cancer patients will respond to immunotherapy?

AI can currently predict immunotherapy response across several cancers better than many existing computational methods, but doctors still need a prospective trial before using those predictions to choose treatment.

This is a big medical problem. Immune-checkpoint inhibitors can produce extraordinary responses in some patients, while many others receive expensive treatment and potentially serious side effects without benefiting. Existing biomarkers such as PD-L1 expression and tumour mutational burden help, but they are imperfect.

COMPASS tries to improve that decision using gene-expression data from tumours. Researchers trained the system on 10,184 tumours across 33 cancer types and represented each tumour through 44 biologically grounded concepts linked to immune activity and the tumour microenvironment.

They then tested it across 1,133 patients from 16 clinical cohorts involving seven cancers and six checkpoint inhibitors.

Across those cohorts, COMPASS beat 22 competing methods. Average accuracy improved by 8.5%, while area under the precision-recall curve improved by 15.7%.

The breadth is what makes the result interesting. Many medical AI models work well on the cancer type, hospital or treatment they were trained around and then weaken elsewhere. COMPASS was designed specifically to transfer across different cancers and therapies.

For now, the model predicts response after looking at existing patient data. The decisive experiment would be to let the prediction change who receives immunotherapy and then compare outcomes with normal biomarker-guided care.

If that works, AI would be doing something clinically much more important than identifying cancer: deciding which expensive, potentially toxic treatment is worth giving to which patient.

Q12Is pathology AI already useful in real hospitals?

Yes, pathology AI is already useful in real hospitals, and the field is simultaneously moving toward foundation models that can handle many pathology tasks with the same underlying system.

A prospective NHS study called Articulate Pro followed 1,613 prostate-biopsy cases across three specialist centres in England. AI assistance was used in 1,049 cases.

When pathologists used AI as a second reader, the system prompted them to revisit cases and changed the original diagnosis or cancer Grade Group in 21 of 386 patients, or 5.4%. Five changes, around 1.3%, were considered potentially capable of affecting clinical management.

Those percentages look modest until we translate them into a normal pathology service. At that rate, an AI second reader would trigger a potentially management-relevant correction roughly once every 77 cases reviewed. The study also found less need for additional immunohistochemistry testing at all three hospitals and a 30.1-hour reduction in average turnaround time under one concurrent-reading workflow.

Meanwhile, pathology models themselves are getting much broader. PRISM2 was trained on 2.3 million whole-slide images and 14 million question-and-answer pairs derived from roughly 700,000 pathology reports. Without building a separate model from scratch for each task, PRISM2 matched or exceeded specialized clinical products on several prostate, breast and lymph-node cancer-detection tests.

Articulate Pro tells us more about current hospital usefulness. PRISM2 tells us more about where the technology is heading. Together, they suggest that pathology AI is moving from narrow classifiers toward general systems that can read slides, connect morphology with clinical language and assist across several parts of the diagnostic process.

100+ new signals every week · 50+ markets · updated daily

Need to know what's hot?We can send you all the signals

Send me the signals Delivered straight to your inbox

Q13Why can a good medical AI still fail in the hospital?

A good medical AI can still fail clinically because the algorithm is often improving a step that was never the real bottleneck.

The LungIMPACT trial shows this unusually well. Across five NHS trusts, researchers analysed 93,326 chest X-rays to test whether AI prioritization could speed up lung-cancer diagnosis.

At one hospital, radiology reporting became meaningfully faster. Median reporting time fell from 47 hours to 34 hours, a reduction of roughly 28%.

The rest of the cancer pathway barely noticed.

Median time from the chest X-ray to CT was 53 days with AI prioritization and 53 days without it. Among patients eventually diagnosed with lung cancer, median time to diagnosis was 44 days with prioritization and 46 days without it, a difference that was not statistically significant. Referral timing, treatment timing and cancer stage also remained broadly similar.

The Kenyan LLM trial showed the same problem from another angle. AI improved documentation and treatment planning, yet short-term treatment failure did not significantly change.

SHAKED exposed a third bottleneck: the doctors themselves. The system's sampled recommendations were almost universally judged appropriate, but usage dropped from 68% to 30% within four weeks.

This is one of the most useful lessons from medical AI right now. Accuracy is only one link in a long chain. If the hospital has a six-week CT backlog, speeding up the X-ray report by half a day changes very little. If an AI recommendation adds friction to an overloaded doctor's shift, good advice can simply go unused.

The next generation of medical AI trials will have to measure whole workflows, not only whether the algorithm was right.

Why technically useful medical AI can have little clinical impact

AI system or trial What improved What barely changed What really limited the benefit
LungIMPACT X-ray reporting speed Time to CT and lung-cancer diagnosis Downstream capacity
Kenyan LLM trial Documentation and care planning 14-day treatment failure Turning better advice into better outcomes
SHAKED Availability of clinically appropriate advice ED length of stay Clinician adoption during busy shifts

Q14Which AI medical breakthroughs are the biggest right now?

Right now, we would put AI-assisted diagnosis and AI drug development ahead of general AI doctors, while live surgical AI is the freshest breakthrough but still the least proven.

The best diagnostic advances have something the chatbot demonstrations still lack: clinical comparisons against normal care. Retina4IRD produced a 21.2-point jump in top-five genetic diagnostic accuracy. The mammography study cut radiologist workload by 63.6% while finding more cancers. LiON found malignant lesions after the initial radiology process had missed them.

Rentosertib deserves a different kind of weight. A diagnostic system can prove useful in hundreds or thousands of patients relatively quickly; a new drug must survive years of failure points. Reaching Phase 3 therefore validates much more of the AI-drug-discovery chain than an early molecule announcement ever could.

Live UCLH neurosurgery currently sits at the opposite extreme. The capability is striking because AI has entered the surgeon's real-time visual loop, but the evidence still consists of the beginning of a clinical trial. The new capability is worth paying attention to. Its effect on surgical outcomes is still unknown.

General medical copilots rank lower for now. They may eventually touch far more patients than any specialist imaging model, but their real-world advantage is surprisingly difficult to prove. Recent studies keep finding useful outputs without equally clear gains in outcomes or workflow.

AI medical breakthroughs ranked by current evidence and potential impact

Current rank Breakthrough Why we rank it here
1 AI improving specialist diagnosis Randomized trials are showing large, measurable gains
2 AI-originated drugs reaching late-stage trials Potential impact is enormous and the validation hurdle is unusually high
3 AI taking over parts of medical imaging workflows Large workload gains are already measurable
4 AI catching findings missed during routine care Direct evidence that AI can act as a diagnostic safety net
5 Real-time AI-assisted surgery Very new capability, but clinical evidence is still tiny
6 General medical AI copilots Impressive reasoning, mixed evidence of actual clinical benefit

Q15Are the latest AI medical breakthroughs really changing medicine?

Yes, the latest AI medical breakthroughs are starting to change medicine, but the change is happening through specific clinical jobs rather than through one universal AI doctor.

We now have randomized evidence that AI can make specialists substantially better at difficult diagnoses. We have prospective evidence that AI can remove much of the routine workload from some screening programs and retrieve cancers overlooked in normal care. AI-originated medicines are reaching late-stage human trials. And surgeons can now receive information extracted from live operative video while they work.

At the same time, some of the largest clinical trials are giving us a useful warning. A highly accurate algorithm can speed up the wrong part of a hospital pathway. A medical LLM can give good recommendations that doctors gradually stop opening. Better documentation can coexist with virtually unchanged patient outcomes.

That mixed evidence makes the breakthroughs that survive it more convincing. Medical AI is no longer impressive simply because a model can answer medical questions or recognize disease on a carefully selected dataset. Those abilities are becoming normal.

The meaningful advances these days are much harder to achieve: finding something a doctor missed, making a specialist measurably more accurate, freeing a large amount of clinical capacity, guiding a surgeon during a live operation, or carrying an AI-generated drug far enough through development that thousands of patients may eventually test whether it really works.

So yes, there are real AI medical breakthroughs happening now. They are narrower than the popular idea of an AI doctor replacing medicine, but they are also much more concrete. AI is beginning to earn a place inside specific parts of diagnosis, drug development and surgery where we can finally measure what it adds.

We track MedTech daily. Want the market signals in your inbox?

Send me the signals
Methodology and sources

The question "what are the latest AI medical breakthroughs?" is harder to answer than it first appears because the word breakthrough is used for very different things: benchmark results, promising research models, prospective clinical trials, systems used on real patients, and drugs progressing through human development.

Rather than relying on intuition, visibility, or a few spectacular announcements, we broke the question into meaningful analytical dimensions, including diagnosis, clinical workflow, physician decision support, surgery, drug discovery, and treatment selection. Within each dimension, we reviewed recent developments and concentrated on the evidence that materially changed what could be said about the field.

We assessed that evidence point by point using several factors: the strength of the underlying study, how close the technology had moved to real patient care, the size of the demonstrated effect, its maturity, and the scale of its potential medical impact. We gave more weight to prospective and randomized clinical evidence than to retrospective benchmarks or simulations, more weight to live clinical use than to demonstrations, and more weight to drugs surviving successive stages of human development than to molecules generated computationally but not yet clinically tested.

When AI improved one part of a medical workflow, we also looked at what happened downstream rather than assuming that a better intermediate metric automatically translated into better care. LungIMPACT, the Kenyan primary-care trial and SHAKED were especially useful here because they separate technically useful AI from measurable improvements in the wider clinical pathway.

Finally, we aggregated the strongest evidence across those dimensions instead of allowing one study, company, or headline to determine the answer. The resulting ranking is not a mechanical score. It is a structured synthesis of where recent evidence is strongest, where AI has moved furthest from technical promise into actual medicine, and where the potential impact is substantial enough to justify the word breakthrough.

In short, we treated "breakthrough" as an evidence question, not a branding term.

Key sources used for this analysis include: the FDA discussion on generative AI-enabled medical devices, JAMA Network Open's review of FDA-authorized radiology AI devices, the LiON prospective liver-malignancy study, the partially autonomous AI mammography trial, the Retina4IRD randomized clinical trial, the DeepRare rare-disease diagnosis study, the multimodal AMIE evaluation, the randomized Kenyan primary-care AI trial, the SHAKED prospective emergency-department study, UCLH's report on live AI-assisted brain surgery, the Rentosertib Phase 2a study, the Rentosertib Phase 3 trial record, the deep-learning antibiotic-discovery study, the COMPASS immunotherapy-response study, the Articulate Pro prospective NHS pathology study, the PRISM2 pathology foundation-model study, and the LungIMPACT randomized lung-cancer pathway trial.

100+ new signals every week · 50+ markets · updated daily

Want to find the next opportunity?We can send you all the signals

Send me the signals Delivered straight to your inbox