27 June 2026 · 6 min read

AI for Science Needs Research Infrastructure

AI for Science

This essay began with a report I read recently from the ARQ Foundation, What we talk about when we talk about AI for science.1 It makes a distinction I have not stopped thinking about, because it lands close to the question I work on every day: what kind of research infrastructure does AI for science actually need? What follows is my attempt to take that thread to its end.

"AI for science" has become one of those phrases that everyone uses and few define. It covers systems that do completely different things, and the confusion is not harmless. When governments decide what to fund, imprecision turns into misallocated money. So it is worth being clear about what we are actually talking about.

Two kinds of system

There are two kinds of system hiding under the one name.

The first models nature. AlphaFold2 predicts protein structures, GraphCast3 forecasts the weather, GNoME4 proposes new materials. These models learn the behaviour of a physical system from data, and once trained they let us see or predict things that were out of reach before. They are cheap to run, and they behave like new instruments. They open doors that were closed.

The second does not model nature. It organises the work of doing research. Built on top of large language models, these systems read the literature, propose hypotheses, plan experiments, call tools, and gather evidence across many steps. They do not produce a new measurement. They coordinate the process of getting one. They are less like an instrument and more like a tireless junior colleague who runs the investigation.

These are not two versions of the same thing. They differ in architecture, in the data they need, in cost, and in purpose. Treating them as one is the first mistake.

The two belong together. Picture them in a single loop. A research system reads the literature and proposes a candidate material. It calls a foundation model to simulate that material's properties. If the simulation looks promising, it proposes an experiment to run in the real world. The foundation model is the precision instrument. The research system decides when and how to use it. Neither is much use alone: a model of nature with nobody to question it, or a tireless investigator with no instruments and no laboratory, reasoning in a vacuum.

The bottleneck is data, and it is layered

What limits both is data, but not the same data, and this is where it gets interesting.

Foundation models need large, specialised datasets: the output of observatories, national laboratories, long measurement campaigns. That data is expensive and scattered, so the difficulty is mostly coordination and standards. Hard, but familiar.

The systems that organise research are short of something stranger. They are trained on the published record of science, and the published record holds results, not the road to them. The judgment, the dead ends, the decision to trust a number or throw it out, the feel for when a clean result is quietly wrong: none of it is written down. It is the part of science learned at the bench, passed on by working next to someone, not the part that reaches a paper.

This is why a bigger model will not, on its own, become a scientist. The recent evidence is sobering and encouraging at once. The strongest current models are more capable at discovery than many expected, and on some tasks they match or beat systems built specifically for the job.5 But they share systematic blind spots on real scientific work, and making them larger returns less and less. Where they improve is not through size. It is through being connected to the tools and checks that scientists already use, and through searching for the right direction rather than recalling a fact.

So the honest version of the claim is this. AI for science is not a smarter model. It is a capable model joined to the instruments, the validators, and the recorded judgment of how science is actually practised. The first two we know how to build. The third we have barely begun to collect.

The laboratory is the apparatus

That third thing is the hardest kind of knowledge to capture. It is tacit, living in habits and hands more than in documents. It is multimodal and messy: instrument readouts, half-failed runs, images and spectra, the texture of a signal that turns out to be noise. And it is hard to coordinate, spread across instruments and laboratories and institutions, expensive to record, and recorded by almost no one, because no career was ever advanced by carefully documenting a failure.

But there is a place where this knowledge is produced all the time, as a byproduct of ordinary work. The laboratory. A research infrastructure that runs self-driving experiments and advanced, multimodal characterisation generates both kinds of missing data at once: the large specialised datasets the foundation models need, and the messy, process-level record the research systems need. The instrument and the experiment are how the tacit becomes recorded. If we want AI that knows how science is done, we have to build the places that do science and capture what happens in them.

This reverses the usual order. We tend to think of research infrastructure as something that will benefit from AI, a consumer of these new tools. It is the other way around. Even the roadmaps toward more capable AI now name experimental data and real-world interaction as a hard limit on further progress. One recent report from a frontier lab put it plainly: AI is not an armchair science.6 The laboratory is not downstream of AI for science. It is on its critical path.

The European opportunity

This is the part that should interest Europe.

Europe is not winning the race for frontier language models. That contest rewarded enormous compute and capital, and it was settled elsewhere. But the layer this argument points to, the physical practice of science, rewards a different set of strengths, and they are ones Europe holds. Europe lost the language-model race, but the physical side of AI is the larger prize, and one it can still win. This is not a consolation. It is the part that turns models into discoveries, and it runs on instruments, domain depth, experimental data, and a willingness to coordinate.

What would it take. Three industries would have to work together, and to do it robustly. The makers of scientific instruments, where Europe has real depth. The software that runs and orchestrates them. And the laboratories and research infrastructures that use them and produce the data. The hard part is not any single one. It is getting them to work as one system, so that an experiment can be planned, run, measured, and learned from without rebuilding the connections each time.

That connecting layer is being defined right now, and not in Europe. The two serious published proposals for it, one for how an automated agent talks to an instrument,7 one for describing a laboratory in a portable way,8 come from China and from Spain. There is an opening here, and an equal risk of the standard being written without us.

Building it does not mean starting over. Europe already has shared materials-data efforts to connect to: NOMAD,9 the OPTIMADE standard,10 and, launched just this month, Materials Commons for Europe,11 a federated infrastructure that links national materials-data hubs across fourteen European countries, from Germany and France to Belgium and Denmark. The Netherlands is not among them. That absence is the concrete first move for public funding here: build the Dutch hub that connects the country's laboratories and their experimental data into this commons, rather than leaving the Netherlands outside it. It is also the right kind of work to fund, because these commons hold mostly computational and structural data, and what they lack is the experimental record that only real instruments produce. The layer worth building is the one that lets that experimental data flow into the commons that already exist, so a measurement made in one laboratory becomes knowledge across the system. Connecting the new to the existing, instead of building another island, is the difference between adding capacity and adding fragmentation.

Europe's weakness was never the quality of its laboratories or its companies. It is that they do not work together. Closing that gap is the task, and it is one private money will not take on. Companies will build the models and capture the value; they will not sustain shared, open infrastructure for the whole research system, or fund the unglamorous work of recording process and failure. Public funding can. It can set the standards, fund the shared data and workflow systems as infrastructure rather than as scattered projects, connect them to what already exists, and build the simple awareness that this data is worth keeping.

For now, AI for science is augmentation, not automation. It becomes powerful when a capable model is joined to instruments and to the tacit, multimodal record of how science is really done. Producing that record is the work of laboratories and research infrastructure. Making it shared, open, and connected is the work of public funding. Europe has the laboratories and the instruments. What it can still build is the connective tissue that turns them into the place where the next way of doing science is invented.

References

  1. B. Özcan, What we talk about when we talk about AI for science, ARQ Foundation, 2026. arq.foundation
  2. J. Jumper et al., "Highly accurate protein structure prediction with AlphaFold," Nature 596 (2021). doi:10.1038/s41586-021-03819-2
  3. R. Lam et al., "Learning skillful medium-range global weather forecasting" (GraphCast), Science 382 (2023). doi:10.1126/science.adi2336
  4. A. Merchant et al., "Scaling deep learning for materials discovery" (GNoME), Nature 624 (2023). doi:10.1038/s41586-023-06735-9
  5. Z. Song et al., "Evaluating Large Language Models in Scientific Discovery," arXiv:2512.15567 (2025). arxiv.org/abs/2512.15567
  6. T. Genewein et al., "From AGI to ASI," Google DeepMind, arXiv:2606.12683 (2026). arxiv.org/abs/2606.12683
  7. L. Zhu et al., "LAP: An Agent-to-Instrument Protocol for Autonomous Science," arXiv:2606.03755 (2026). arxiv.org/abs/2606.03755
  8. S. Pablo-García et al., "A foundational representation for an orchestrated lab," Device (Cell Press), 2026. doi:10.1016/j.device.2026.101202
  9. NOMAD, the FAIR materials-data infrastructure. nomad-lab.eu
  10. OPTIMADE, an API standard for interoperable materials databases. optimade.org
  11. Materials Commons for Europe, a Horizon Europe Innovation Action (HORIZON-CL4-INDUSTRY-2025-01-MATERIALS-45), launched June 2026. materialscommons.eu
← Back to Blog