Popper-consensus · does it know when it's wrong?

Does compile-time knowledge organization generalize?

Popper showed a knowledge-graph compiler beats its controls on five hand-curatedbenchmark literatures. The obvious objection: those corpora were built by the same person who built the compiler. So we hold the entire methodology fixed — same schema, same scoring rubric, same grounding audit, same four-way control structure — and change only the corpus: from curated benchmarks to the Open Research Knowledge Graph (6.3M triples, 65,689 papers, 8,420 research problems). The only independent variable is the corpus.

compiler
0.99
mean composite (ORKG)
llm-only
0.00
mean composite (ORKG)
keyword
0.59
mean composite (ORKG)
random
0.00
mean composite (ORKG)

The claim under test: structured knowledge compilation produces higher-quality, grounded, falsifiable hypotheses than unstructured generation — and that it holds on a large, heterogeneous corpus the author did not curate. If the compiler bar is not clearly above the three controls here, the claim fails to generalize, and that is the result.

20
ORKG domains
7
scientific disciplines
7
adversarial false theories
191
compiled hypotheses

Per-discipline generalization

Does the compiler hold across fields, or do some disciplines expose weaknesses in the schema?

biomedicine
3 domain(s)
composite0.98
grounding1.00
rediscovery0.00
chemistry
3 domain(s)
composite0.99
grounding1.00
rediscovery0.07
computer science
3 domain(s)
composite0.99
grounding1.00
rediscovery0.00
environmental science
3 domain(s)
composite0.98
grounding1.00
rediscovery0.10
materials science
3 domain(s)
composite0.99
grounding1.00
rediscovery0.00
neuroscience
2 domain(s)
composite0.99
grounding1.00
rediscovery0.00
physics
3 domain(s)
composite0.98
grounding1.00
rediscovery0.00

3 hypotheses per method per domain · audit off · compiled locally on a 35B model.