Bounded substrate comparison
Benchmark performance and biological plausibility
How should task utility and resemblance to biological mechanisms coexist?
engineered evaluation
Task performance
Valid claim: Meets this metric on this benchmark.
model-to-biology comparison
Biological plausibility
Valid claim: Matches these selected biological observations under this test.
Comparable axes
- Declared model
- Observed outputs
- Uncertainty and alternatives
Non-equivalences
- High accuracy does not imply brain likeness.
- Biological resemblance does not imply engineering utility.
Comparison procedure
- 1.Score task independently.
- 2.Test biological predictions independently.
- 3.Publish a two-axis result.
Prohibited inference
Do not collapse benchmark rank and biological plausibility into one intelligence score or use success on either axis to certify the other.