Blog

Abundant Data, Scarce Knowledge: Time to Attack the Quantification of Knowledge

After more than two decades of data arms race, biomedicine received a cold verdict: data explosion, knowledge poverty. Knowledge has long lived in three forms—travelogues (literature reviews), natural history museums (databases), and maps (computable knowledge)—and biomedicine lacks precisely the third. This essay argues from the history of science that a discipline matures when its core knowledge turns from narrative into computable objects: astronomy has star catalogs, chemistry has the periodic table, biomedicine has yet to build its coordinate system. It is time to attack the quantification of knowledge: parameterize, structure, and close the loop—so that a worldview finally acquires a computable carrier.

熊江辉 · 2026-08-17
2

In August 2026, Nature Reviews Drug Discovery published a review jointly signed by sixteen scholars, systematically reassessing a decade of AI applications in drug discovery.

Spanning more than twenty pages, its verdict can be compressed into four words:

> Data explosion, knowledge poverty.

More than two decades have passed since the draft human genome was published; sequencing costs have fallen off a cliff, and multi-omics data have grown exponentially. Yet the review's judgment remains sober: mere acquisition produces only data; knowledge is what can dissect biological processes at the level of function and support real decisions.

That sentence deserves a pause from the entire industry.

Because it points to a fact of stark contrast:

> Sequencers are refreshed every few months, yet the container of knowledge has not changed in over three hundred years—it is still that journal article, merely recast from movable type into PDF.

3

1. Three Forms of Knowledge: Travelogues, Natural History Museums, and Maps

If we roughly classify the biomedical knowledge humanity has accumulated, it falls into three forms.

The first form: the travelogue.

The literature review is the archetypal travelogue. A senior researcher spends two years digesting several hundred papers in a field, then composes a review of rigorous logic and elegant prose. It tells you where the mountains rise, where the rivers run, and which roads are passable.

Take aging research. A 2023 Cell review expanded the Aging Hallmarks to twelve, setting out the mechanism, evidence, and interrelations of each hallmark with impeccable clarity. It is an excellent travelogue.

But the travelogue carries an inherent limitation:

> It can be read, but not computed.

Twelve aging hallmarks—but which hallmark, in which organ, with what weight, through which pathway, contributes to which disease? The text yields no numbers. You can cite it and debate it, but you cannot place it inside an equation and compute with it.

The second form: the natural history museum.

Databases take the form of the natural history museum. UniProt stores over a hundred million protein sequences; the GWAS Catalog holds hundreds of thousands of genetic associations. Every "specimen" is neatly labeled, clearly indexed, and ready for retrieval.

But the natural history museum's limitation is equally plain:

> Querying is not reasoning.

You can ask "what known associations does this gene have," but it will never tell you "how will the entire network respond once this module is perturbed." Specimens are static; there are no roads between them.

The third form: the map.

The map is the third form of knowledge—and the rarest.

The fundamental difference between a map and the travelogue or the museum is this: on a map, every object has coordinates, the objects are joined by a road network, and the road network supports route planning.

> In other words: the travelogue describes terrain; the map supports simulation.

With a quantified map in hand you can ask: starting from here, which road should I take, at what cost, and where does it lead? This question cannot even be posed inside a travelogue or a museum—the former lack coordinates, the latter lacks roads.

What biomedicine lacks most is precisely this third form.

4

2. Quantification Is the Critical Leap for Knowledge

Looking back across the history of science, one pattern recurs:

> A discipline matures when its core knowledge turns from narrative into computable objects.

Astronomy moved from astrological narrative to star catalogs and the equations of universal gravitation—and only then could orbits be forecast. Chemistry moved from the alchemists' recipe narratives to the periodic table and quantitative reactions—and only then was a synthetic industry possible. Physics moved from the natural philosophers' investigation of things to differential equations—and only then was engineering born.

Every such leap was accompanied by a change in the container of knowledge: from the article to the coordinate system.

And biomedicine?

It commands the largest data output and the most prolific literature system of any discipline, yet its core knowledge still exists overwhelmingly in narrative form. Reviews grow ever longer and mechanisms are drawn ever finer, but the moment one asks a quantitative, computable question—"through this module, how much change will this intervention produce in this person?"—the narrative falls silent.

This is the precise meaning of "data explosion, knowledge poverty": not that knowledge is too scarce, but that it is not computable enough.

Making knowledge computable requires three moves:

- Parameterization: every knowledge object acquires measurable parameters. An aging hallmark is no longer merely "cellular senescence drives inflammation," but a quantity with a reading, a unit, and a normal range.

- Structuring: knowledge objects acquire composable relations. How modules connect, how they are redundant with one another, how they compensate for one another—these are written into network structure, not into subordinate clauses.

- Closing the loop: knowledge can be tested, corrected, and iterated. When a prediction fails, the parameters change with it; when new data arrive, the boundaries are adjusted in step.

Parameterization makes knowledge measurable; structuring makes it computable; closing the loop makes it evolvable.

Only when all three are in place does knowledge turn from something you read into something you run.

5

3. A Worldview Needs a Carrier

One step deeper: why is the quantification of knowledge so important?

Because a worldview needs a carrier.

Every serious researcher carries a biomedical worldview in their head—an entire architecture of belief about how the body operates, how disease arises, and how interventions take effect. But as long as that worldview lives only in the head, it suffers three fatal defects: it cannot be shared, it cannot be computed, and it cannot resist forgetting.

Your worldview differs from mine, so we can argue—but we can never audit each other.

> A worldview must not live only in the head; it must live in a coordinate system.

A quantified knowledge map is the public carrier of a worldview. It converts "I believe inflammation drives aging" into "what is the coupling coefficient between the inflammation module and the aging module, what is the level of evidence, and in which populations has it been validated?"

Only once it is loaded into a coordinate system can a worldview do what no in-the-head worldview can ever do:

> Simulate.

This is precisely the dividing line between a qualitative worldview and a world model (as I have discussed in my world-model series): a qualitative worldview can only describe terrain; only a quantified worldview can answer "if this action is taken, how will the system change?" Each of the five elements of a world model—State, Action, Transition, Objective, and Feedback—requires a coordinate system on which to land.

Without quantified knowledge, there is no simulable world model; without a simulable world model, the word "precision" remains forever a figure of rhetoric.

6

4. A Roadmap for the Attack

A judgment demands a strategy. The attack on the quantification of knowledge proceeds, roughly, in four steps.

Step one: draw the module boundaries.

Quantifying knowledge does not mean digitizing the entire literature; it begins by answering "into which computable knowledge objects should the human system be partitioned?" In our research with the Future Laboratory of Tsinghua University, we built a human-scale representation composed of 332 modules—covering the expanded aging hallmarks, organ systems, immunity, and nutrition and food-based interventions, and also incorporating Traditional Chinese Medicine (TCM) syndrome proxies and medicine-food homology targets. A literature-derived knowledge network draws the module boundaries: this is the knowledge-driven half.

Step two: learn the module parameters.

Multi-omics data learn the parameters of each module: this is the data-driven half. Two legs walking together, neither dispensable—knowledge without data leaves the map empty; data without knowledge leaves the coordinates without semantics.

Step three: hold the map to the test.

Once quantified, a map must withstand comparison. We placed this 332-module map under systematic evaluation in a drug repurposing setting—1,916 DrugBank small molecules, five chronic-disease tasks (with exploratory extrapolation to 23 disease categories)—first forge the ruler, then use the ruler to measure the map.

Step four: report the results faithfully, the unfavorable ones included.

One finding deserves to be singled out: across the 23 disease categories, the full-module map was strictly optimal in only nine. There is no universally optimal biological map; each disease favors its own best-fitting combination of representations.

This "negative result" is itself exactly the kind of knowledge that only quantification can deliver: in the era of narrative, "which representation is better" was not even a rigorously answerable question. Only when the map acquires coordinates and scores do you discover that the value of a map is scenario-dependent—and that "which map to choose" is itself a decision variable.

7

Conclusion

Data is ore, knowledge is smelting, and the quantified knowledge map is the step that turns ore into navigation.

Astronomy waited for its star catalogs; chemistry waited for its periodic table. Remember: when Mendeleev drew up his table, several cells were still empty—a coordinate system need not wait for all the facts to arrive before it is drawn; it reserves seats for the unknown in advance. Biomedicine already has data in surplus; what remains owed is an attack on the quantification of knowledge:

> Turn the knowledge inside reviews into coordinates, the specimens inside databases into road networks, and the worldview carried in our heads into a carrier that is measurable, computable, iterable, and auditable.

The more precisely a discipline can pose its questions, the more steadily it advances. This attack deserves to begin now.

References

1. Bender A, Thomas MC, Scannell JW, et al. Artificial intelligence in drug discovery — what it is, where we stand and the path forward. Nature Reviews Drug Discovery. 2026. DOI: 10.1038/s41573-026-01496-2.

2. López-Otín C, Blasco MA, Partridge L, et al. Hallmarks of aging: An expanding universe. Cell. 2023;186(2):243-278.

3. Xiong J, Xia Q. Toward a Self-Learning AI Agent for Drug Repurposing: Building Human-Scale Representations for Virtual Patients. Preprints.org. 2026. DOI: 10.20944/preprints202608.0998.v1.

4. Xiong J. World Models for Biomedicine: A Steerability Framework. Preprints.org. 2026. DOI: 10.20944/preprints202605.0366.v1.

Related Reading

Blog

From Virtual Cell to Virtual Patient: The Missing Layer in Between

The virtual cell is having its moment: a Cell paper by forty-plus authors proposed building an AI Virtual Cell with multi-scale foundation models. But a patient is not a bigger cell—from cell to human body lie multiple emergent transitions, where pathway redundancy and patient heterogeneity keep making 'works in vitro' end at 'fails in vivo'. This essay argues that the path from virtual cell to virtual patient is missing not a bigger model, but a map at another scale—the 'molecule → module → human' conversion layer—and that virtual patients require two things virtual cells cannot offer: the medical semantics of Action (will this person respond to this intervention?) and a re-test feedback loop (predict → intervene → re-test → correct). The essay closes with the economics: the Phase II valley of death, ~$0.9 billion per approved drug, and biomarker stratification halving costs—the key is not under the streetlight.

Essay

Life as an Adaptive Capability Ensemble: Proposing and Demonstrating the First Principle

Proposing the central axiom: life is an adaptive capability ensemble. From this axiom, deducing the five core functional modules that any living system must possess.

Essay

Is a Whole Piece of the Puzzle Missing from Modern Medicine's Knowledge Paradigm?

Medicine has two knowledge paradigms: objective description of external entities, and response mapping of system states. Modern medicine excels at the first but has a systematic blind spot in the second.

Back to Blog