Prospective evaluation and falsifiability protocol

Ordinary-baseline comparison protocol

Require an astrology model to outperform a declared non-astrological comparator on identical tasks.

Required inputs

Locked astrology model

Ordinary baseline

Identical eligible tasks

Scoring rule

Minimum sample

Ordered execution

Work the protocol

  1. 1

    Choose the baseline before outcomes.

  2. 2

    Give both models identical non-outcome information.

  3. 3

    Score all tasks including abstentions.

  4. 4

    Calculate paired differences.

  5. 5

    Report the baseline even when both models fail.

Outputs

  • Paired score table
  • Coverage and abstention rates
  • Baseline comparison verdict

Done only when

  • Task sets are identical.
  • The comparison is prospective.
  • Uncertainty and null results are reported.

Fail closed

Refusal conditions

  • Only astrology successes are sampled.
  • The baseline gets less information.
  • The comparator is changed after scoring.