A system running 10,000 autonomous AI agents in parallel for 88 hours just produced an analytical proof for a problem that had resisted mathematicians for well over a century. The problem was the Navier-Stokes equations, one of seven Millennium Prize Problems in mathematics.
That result is not a curiosity. It is a signal about what AI can now accomplish inside genuinely hard scientific domains, the kind where being approximately correct is worthless and only rigour counts.
Over the past year, artificial intelligence has been quietly shifting from a tool that generates content into something closer to an engine for scientific inquiry itself: forming hypotheses, running simulations, and validating results across pharmaceuticals, fusion energy, and materials science. Institutions including UBS Wealth Management, Goldman Sachs, and Morgan Stanley are now framing this as a long-duration thesis with multi-trillion-dollar addressable markets by 2030. This is not a near-term earnings story. It is a revaluation of AI’s commercial ceiling.
This analysis maps the concrete evidence behind that thesis, so you can judge which parts of the AI value chain are genuinely de-risked, which remain speculative, and where the structural bottlenecks sit before any capital is committed.
OpenAI’s Navier-Stokes proof and what it reveals about multi-agent scientific reasoning
On 8 September 2026, OpenAI published a result that pushed past the boundary of what AI had previously demonstrated in formal mathematics. Working from a fluid initially at rest under a smooth, finite-energy external force, the system located conditions under which fluid forces grow without bound, triggering a finite-time singularity, a catastrophic mathematical breakdown in the equations.
The scale of the effort matters. A multi-agent system of 10,000 autonomous agents ran simultaneously over an 88-hour generation phase to construct the proof.
Energy infrastructure bottlenecks represent the most concrete near-term constraint on the multi-agent compute scale demonstrated here: running 10,000 simultaneous agents across 88 hours at the power density of modern AI accelerator racks sits inside a demand environment where the IEA projects data centre and AI electricity consumption will exceed 1,000 TWh by 2026.
Then came the step that separates this from every prior AI mathematical claim. The proof was formally verified in Lean 4, a machine-checked mathematical validation platform, by a model referred to as GPT-6 Astra, over roughly 17 hours.
The Lean 4 verification is the detail investors should focus on. Formal machine-checking transforms the output from a plausible AI-generated argument into a proof that satisfies the highest standard of mathematical rigour. Multi-agent systems can now produce results that meet that bar, not merely approximate it.
Navier-Stokes sits among the seven Millennium Prize Problems, and OpenAI has stated it does not intend to claim the prize. The downstream applications named are more telling than the accolade: aviation aerodynamics and, notably, nuclear fusion energy.
Why the architecture matters more than the result
The specific proof is impressive. The architecture behind it is the investable part.
The three-phase structure looks like this:
- Multi-agent generation: 10,000 agents working in parallel for 88 hours to construct the analytical proof
- Formal verification: Lean 4 machine-checking by GPT-6 Astra over approximately 17 hours, confirming validity
- Publication and application: released 8 September 2026, with aerodynamics and fusion cited as downstream fields
This pattern, large-scale multi-agent reasoning paired with formal verification, is what institutions are watching across every scientific domain. It is the same “Discovery Engine” framework being applied to the three commercial sectors examined below. Understanding it is the prerequisite for understanding the rest of the thesis.
When big ASX news breaks, our subscribers know first
Drug discovery pipelines: 170-plus programs in development, zero approvals, and what the gap means
The pipeline numbers are the strongest part of the pharmaceutical case. In late 2023, roughly two dozen AI-designed drug programs existed in clinical development. By mid-2026, tracking sources counted more than 170. That is not incremental progress; it is a change in the production model of early drug discovery.
The exact figure depends on the source and the counting method.
| Source and date | AI drug programs in clinical development | Stage breakdown |
|---|---|---|
| IntuitionLabs, August 2026 | Over 173 globally | Not specified |
| Axis Intelligence, June 2026 | 200-plus clinical-stage candidates | Approx. 94 Phase I, 56 Phase II, 15 Phase III |
| ASCO, December 2025 | 117 AI-enabled assets across 63 companies | Interventional human trials |
The individual case studies make the progress concrete:
- Insilico Medicine: ISM001-055 (Rentosertib), the first end-to-end generative-AI-discovered target and molecule to reach Phase II, has now advanced to Phase III for idiopathic pulmonary fibrosis (a progressive lung-scarring disease).
- Recursion Pharmaceuticals: reported positive preliminary Phase II efficacy for REC-4881, and has secured over $525 million in upfront and milestone payments from Roche/Genentech and Sanofi partnerships.
- Absci: became clinical-stage in 2025 with ABS-101 entering Phase I, its first program to reach human trials.
Impressive as those milestones are, there is a caveat of equal weight. As of September 2026, no novel AI-discovered drug has received full regulatory approval anywhere. Not one.
The clinical AI commercialisation gap is measurable well below the pipeline headline numbers: as of September 2024, clinical AI showed only 6.8% commercial maturity across categories, a figure that sits alongside the zero-approval count for AI-discovered drugs and reinforces why the gap between discovery volume and revenue realisation is the central tension in the pharmaceutical thesis.
Why the regulatory gap is the key variable for investors
The gap between 170-plus programs and zero approvals is not a failure signal. It is a valuation calibration tool. It tells you this sector carries early-stage biotech risk, not mature pharma risk, and positions should be priced accordingly.
The regulatory window has not shortened. The European Medicines Agency’s 2026 reflection paper now requires formal risk analyses for AI and machine-learning applications, adding a compliance layer rather than removing one. Industry reviews note that AI has yet to demonstrate a consistent reduction in the 6-to-8-year clinical trial timeline.
Here is the mechanism to hold onto. AI compresses the pre-clinical discovery phase, the part where molecules are designed and screened, but it has not changed the clinical validation window, where drugs are tested in humans over years. That means the investable action today sits in enabling infrastructure and in milestone-based partnership structures like Recursion’s, not in direct exposure to approval outcomes that remain years away.
Fusion plasma control and materials screening: the infrastructure thesis behind both sectors
Fusion and materials science look like different stories. For investors, they share a single logic: AI’s commercial value in both is primarily in reducing experimental cost and failure rate, not in generating near-term revenue. Both are long-horizon positions where the science is being de-risked well ahead of the economics.
The milestones separate cleanly once that frame is set. Ranked by deployment maturity, from experimental to fully integrated:
- DeepMind at DIII-D (February 2024): A Nature paper documented deep reinforcement learning deployed to avoid tearing-mode instability, a foundational demonstration building on earlier DeepMind and EPFL collaboration.
- China’s HL-3 tokamak (October 2025): An LSTM-based AI controller was fully integrated into the reactor’s live plasma control system, achieving stable operation and zero-shot generalisation under unfamiliar conditions.
- PACMAN framework (September 2026): Developed by Princeton University and the Princeton Plasma Physics Laboratory for the DIII-D tokamak, running fast control loops every approximately 20 milliseconds, forecasting plasma disruptions roughly 200 milliseconds in advance and autonomously adjusting heating and magnetic parameters.
- Microsoft and DeepMind materials screening: platform-scale discovery at a magnitude no human process can match.
On the materials side, the numbers reframe what discovery even means. Microsoft‘s Azure Quantum team and the Pacific Northwest National Laboratory screened over 32 million candidate materials, then synthesised a working solid-state electrolyte prototype in the Na-Li-Y-Cl family that reduced lithium dependency. Separately, DeepMind’s GNoME model discovered 2.2 million stable crystal structures, adding roughly 381,000 new materials to the convex hull of known possibilities.
AI screened 32 million candidate materials to produce a single working battery electrolyte. That scale of search, running faster than any laboratory could physically test, is the advantage investors are underwriting.
The 200-millisecond disruption warning and the 32-million-candidate screen point to the same conclusion. AI now operates faster and at greater scale than any human experimental process in these fields. That shifts the bottleneck away from discovery entirely and onto commercialisation. Fusion still faces enormous capital expenditure, long construction timelines, and grid-integration complexity. Materials still face synthesis and manufacturability limits. The investable thesis today is in the AI platforms enabling the discovery, not the downstream applications that remain years from scale.
Cross-cutting risks that apply to every sector: IP constraints, commercialisation lag, and the approval gap
After three sectors of compelling evidence, honesty requires the harder accounting. Several systematic risks constrain pharma, fusion, and materials simultaneously, and they leave the sophisticated reader neither bull nor bear.
The most underappreciated risk is intellectual property. Both the US Patent and Trademark Office and the UK Supreme Court, in Thaler v Comptroller of Patents, have ruled that AI systems cannot be named as inventors. Discoveries generated without a significant human contribution to conception may not be patentable at all.
The IP ruling is the risk most investors are not pricing. If a meaningful share of AI-generated discoveries cannot be patented without demonstrable human inventive input, the competitive moats of AI-native biotech and materials firms could be far narrower than their pipelines suggest.
The three cross-cutting risk categories every position must account for:
- IP and patent law: AI cannot be an inventor; discoveries lacking significant human conception may be unpatentable across jurisdictions
- Commercialisation lag: discovery accelerates faster than the ability to turn discoveries into products
- Regulatory approval timelines: the clinical and safety validation windows remain intact regardless of discovery speed
That commercialisation lag is measurable. Economic studies of AlphaFold’s impact found that AI accelerates basic research output by 15-40%, while the rate at which those outputs reach experimental commercialisation lags well behind.
Governance quality as a valuation input has moved from a soft consideration to a live pricing variable, with Morgan Stanley explicitly modelling a heavy-oversight scenario in which compliance costs compress AI technology multiples and a light-touch scenario in which regulatory risk remains structurally under-priced in current forward estimates.
AI accelerates basic research output by 15 to 40 percent, but the rate of experimental commercialisation lags significantly behind. That contrast is the central tension the thesis has to hold.
The synthesis is what matters for positioning. Across all three sectors, the 6-to-8-year clinical timeline holds, IP protection is uncertain, and commercialisation trails discovery. The value chain position with the clearest near-term monetisation is the AI platform and compute infrastructure layer, not the downstream application layer. That is why UBS, Goldman Sachs, and Morgan Stanley frame this as a long-duration, high-risk thesis rather than a near-term trade, even while pointing to those multi-trillion-dollar 2030 markets.
What the evidence actually supports for investors positioning across the AI discovery value chain
The verdict is not a single view. It is a tiered one, because the evidence supports different levels of conviction at different points on the value chain.
| Value chain position | Representative assets or sectors | Time horizon | Key risk |
|---|---|---|---|
| Compute and infrastructure | Chips, cloud compute, power supply for AI workloads | Near-term, highest conviction | Capacity and demand cyclicality |
| AI platform and discovery software | Discovery Engine frameworks, multi-agent platforms | Medium-term, evidence-supported | Monetisation model still maturing |
| Downstream application assets | AI-native biotech, fusion, materials firms | Long-duration, speculative premium | Approval gap and IP defensibility |
What the evidence does not yet support is worth stating plainly. There is no approved AI-discovered drug, no commercial fusion plant, and no room-temperature superconductor in production. The thesis is validated in the science layer. It is not yet validated in the revenue layer.
The AI capital cycle provides the macro backdrop against which every discovery-layer thesis must be assessed: hyperscalers are projected to consume approximately 94% of operating cash flow on AI infrastructure in 2026, compared with a historical average of roughly 40%, a compression that raises the monetisation bar for any platform or application-layer position dependent on that spending.
That distinction sets the positioning logic. The infrastructure layer offers the most defensible near-term exposure. The application layer carries a speculative premium that current IP and regulatory reality may not yet justify.
The re-rating event to watch is specific. The first full regulatory approval of a novel AI-designed drug would materially re-rate the entire thesis, and Insilico Medicine‘s Rentosertib, the leading generative-AI-discovered molecule in Phase III, is the most plausible candidate to deliver it, potentially in the 2027-2028 window.
Three milestones to monitor as leading indicators:
- The Phase III outcome for Insilico’s Rentosertib
- The first live commercial-scale fusion AI control deployment
- The first patented AI-assisted materials discovery to reach production
This article is for informational purposes only and should not be considered financial advice. Investors should conduct their own research and consult with financial professionals before making investment decisions. Past performance does not guarantee future results, and financial projections are subject to market conditions and various risk factors.
Where the AI scientific discovery thesis stands in September 2026
The distance travelled in a single year is the honest measure of this thesis. The science layer is now demonstrably de-risked across mathematics, drug discovery, plasma control, and materials screening. The most recent formal proof point, OpenAI’s machine-verified Navier-Stokes result, was published on 8 September 2026, days before this writing.
AI-designed drug programs went from roughly two dozen in late 2023 to more than 170 by mid-2026. The discovery engine is real and scaling. The commercial engine is still being built.
That acceleration is the most honest indicator of trajectory available. It tells you the discovery capability exists and compounds. It does not tell you when revenue arrives, because the commercial layer, approvals, plants, patented products, remains largely unvalidated.
The next 12 to 18 months contain binary outcomes worth watching closely:
- Insilico’s Rentosertib Phase III readout, the leading approval candidate
- Finalisation of the EMA’s AI and machine-learning regulatory framework
- A fusion AI deployment milestone at a commercial-scale facility
These are speculative and subject to change based on market and company developments. What the reader leaves with is a calibrated position rather than a binary one: the evidence is substantial, the risks are specific and named, and the catalysts are identifiable. That is the difference between informed exposure and speculative momentum.

