Popper-consensus · does it know when it's wrong?

Method

Popper treats a hypothesis as a compiled artifact. A traditional compiler turns source code into an executable that must pass tests. Popper turns literature into a hypothesis that must pass validation.

Literature → Knowledge Compiler → Hypothesis → Validation

The falsifiable hypothesis schema

A hypothesis is admitted only if it is structurally falsifiable. Missing an independent variable, a falsification condition, a quantitative prediction, or a source id is not a stylistic lapse — it is a compile error. Required fields:

The four scores

Composite is the geometric mean of the applicable scores. A zero in any dimension tanks the composite — by design. A beautifully written but ungrounded hypothesis must not score well.

The experiment: only the corpus changed

This iteration holds the entire Popper pipeline fixed (the popper/ package is byte-for-byte identical, md5-verified) and changes only the input corpus — from five curated benchmark literatures to the Open Research Knowledge Graph: 6.3M triples, 65,689 papers, 8,420 research problems. The single independent variable is the corpus.

Temporal rediscovery: for each ORKG research problem we sort papers by year, hold out a later "discovery" paper, and give the compiler only strictly-earlier prior literature(the temporal wall). Successes and failures are both reported; a failure is evidence the wall holds and no future information is leaking.

Four-way controls: compiler vs. llm-only (no corpus) vs. keyword co-occurrence vs. random traversal — on the same input. Component ablation: remove one compiler component at a time (graph reasoning, grounding verification, gap detection, falsifiability eval, literature synthesis) and measure the degradation, to prove the architecture is load-bearing.

Adversarial (attractive nonsense): seven historical dead ends — cold fusion, phlogiston, N-rays, luminiferous aether, caloric, Lamarckism, miasma. A trustworthy system must notreconstruct these as grounded discoveries. Grounding audit: every evidence item is tiered fully / partially / unsupported, with hallucinated out-of-corpus ids counted.

Local-first

Every generation, audit, and score is produced by a 35B model running on the author's own hardware via an OpenAI-compatible endpoint — no cloud APIs, no keys. The whole point is a scientific instrument you can own, inspect, and rerun.