Evaluation and governance

Benchmarking, energy, and task equivalence

Compare systems only after aligning tasks, correctness, boundaries, amortization, and excluded costs.

hybridestablished research

Working definition

A credible benchmark declares the task, dataset, preprocessing, accuracy or quality constraint, latency definition, measurement instrument, system boundary, idle power, host and interface costs, training or adaptation, repetitions, and uncertainty. Chip energy, wall-plug energy, biological metabolic cost, and laboratory support are different quantities and cannot share one unlabeled efficiency ranking.

Mechanism

  • Freeze task and correctness criteria.
  • Declare measurement and amortization boundaries.
  • Report paired performance, resource, and uncertainty metrics.

Measurements

  • Task quality
  • Latency and throughput
  • Energy, materials, labor, and support costs

Reproducibility controls

  • Version hardware, software, firmware, and analysis code.
  • Declare dataset, preprocessing, random seeds, and measurement boundary.
  • Report repeated runs, variation, exclusions, and failed trials.

Limits and failure modes

  • No benchmark covers general usefulness.
  • Cross-substrate totals require explicit accounting models.

Mathematical connection

Formal structure without substrate erasure

Predeclared task scoring

Score probabilistic or categorical outputs under a rule selected before benchmark results are known.

Inputs

  • Frozen predictions
  • Observed labels
  • Declared score

Outputs

  • Comparable task score
  • Uncertainty interval
  • Baseline difference

Limit: A proper score evaluates the declared forecasts; it does not make mismatched tasks, energy boundaries, or substrates equivalent.

Technical and governance sources

  1. [1]NeuroBench: Advancing Neuromorphic Computing Through Collaborative, Fair and Representative Benchmarking · National Institute of Standards and Technology

    Establishes: A community framework separating algorithm and system tracks and defining task, correctness, efficiency, and reporting procedures intended to make neuromorphic results more comparable and reproducible.

    Boundary: A benchmark ranks submitted systems on declared tasks and metrics. It does not prove general intelligence, biological equivalence, safety, usefulness outside the benchmark, or superiority under unreported host and data costs.

  2. [2]Taking Neuromorphic Computing to the Next Level with Loihi 2 · Intel Labs

    Establishes: An official description of the Loihi 2 research chip, its programmable neuron models, event-based communication, on-chip learning support, and the Lava software framework used to construct neuromorphic applications.

    Boundary: This is a vendor technical brief about a research platform. Performance and efficiency results remain workload-, configuration-, measurement-boundary-, and comparison-dependent and do not establish equivalence to biological intelligence.

  3. [3]Molecular computation of solutions to combinatorial problems · Science

    Establishes: A foundational experiment using molecular biology operations and DNA strands to encode and recover a solution to one small directed Hamiltonian-path instance.

    Boundary: The experiment demonstrates a bounded molecular computation. It does not establish practical general-purpose DNA computing, favorable end-to-end energy or latency, autonomous operation, or scalability beyond the reported instance.

Related concepts