Does compile-time knowledge organization generalize?
Popper showed a knowledge-graph compiler beats its controls on five hand-curatedbenchmark literatures. The obvious objection: those corpora were built by the same person who built the compiler. So we hold the entire methodology fixed — same schema, same scoring rubric, same grounding audit, same four-way control structure — and change only the corpus: from curated benchmarks to the Open Research Knowledge Graph (6.3M triples, 65,689 papers, 8,420 research problems). The only independent variable is the corpus.
The claim under test: structured knowledge compilation produces higher-quality, grounded, falsifiable hypotheses than unstructured generation — and that it holds on a large, heterogeneous corpus the author did not curate. If the compiler bar is not clearly above the three controls here, the claim fails to generalize, and that is the result.
Per-discipline generalization
Does the compiler hold across fields, or do some disciplines expose weaknesses in the schema?
3 hypotheses per method per domain · audit off · compiled locally on a 35B model.